12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Peer Review Report For: Evaluation of automated assessments of systematic review adherence to the PRISMA 2020 statement: study protocol [version 1; peer review: 1 approved with reservations]

Matthew J. Page, Evan Mayo-Wilson, Minyan Zeng, David PQ Clark, Daniel G. Hamilton, Phi-Yen Nguyen et al. · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.5256/f1000research.198806.r482712

Methodology & findings

Study design

Diagnostic test accuracy analysis framework (validation study).

Sample

N = 200, 4 groups

Primary method

Python (version 3.14.3) will be used for all analyses. Methods include: (1) Diagnostic test accuracy analysis framework with consensus human response as reference standard; (2) Calculation of frequencies and percentages for systematic review characteristics; (3) Median and interquartile range for token counts, costs, and generation times; (4) Performance metrics: percentage agreement, Gwet's Agreement Coefficient (AC), sensitivity, specificity, positive predictive value, negative predictive value, F1 score; (5) Bootstrap percentile method with 5000 replications for 95% confidence intervals; (6) Clustering of observations within systematic reviews by resampling entire reviews; (7) Systematic comparison of supporting evidence extracted by LLMs and humans

Main result

This is a study protocol, not a completed empirical study. However, the authors state: "We will determine which questions in our comprehensive tool for assessing adherence to PRISMA 2020 can be accurately automated by LLMs. This knowledge will help inform which questions need the most human oversight by meta-researchers, peer reviewers and other interest holders seeking to assess adherence." The study aims to "evaluate the performance of each LLM assessment against the reference standard by calculating percentage agreement, Gwet's Agreement Coefficient, sensitivity, specificity, positive predictive value, negative predictive value, and the F1 score."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude: "Our study will determine which questions in the PRISMA-Check tool can be accurately automated by LLMs. This knowledge will help inform which questions need the most human oversight by meta-researchers, peer reviewers and other interest holders seeking to assess adherence." They further state: "We recognise that the number of questions in our adherence tool is large and that authors and editors of a systematic review might not want a report providing many Yes/No judgements. Future research could build on our planned work to identify how best to consolidate the adherence assessments and to develop tools that provide constructive feedback to authors for improving their systematic reviews."

Risk of bias

Selection bias: Systematic reviews limited to PubMed Central Open Access Subset from September 2025, excluding paywalled reviews; Language bias: Only English-language reviews included; Indexing bias: Limited to one indexing month (September 2025); Scope bias: Limited to intervention reviews on human health; excludes diagnostic accuracy, prognostic, and scoping reviews; Reference standard bias: Human consensus treated as gold standard despite acknowledged imperfections; Exclusion bias: 76 of 171 elements excluded from PRISMA-Check assessment; reviewer notes 'Difficulty is exactly where users will want LLM help, so cutting those out makes the validation easier on the models'; Prompt development bias: Prompts developed on training set; reviewer notes 'The authors should also state plainly that no test-set review will be opened, viewed, or skimmed by anyone during prompt development'; LLM selection bias: Limited to specific LLM versions available in April 2026; 'Preview' versions subject to change; Operator bias: One investigator (DPQC) screens all 200 abstracts and conducts all 100 full-text assessments in test set; Selection bias: Limited to PubMed Central Open Access Subset, excluding paywalled reviews; Language bias: English-language reviews only; Temporal bias: Systematic reviews indexed in one month (September 2025); Population bias: Intervention reviews on human health only; excludes diagnostic accuracy, prognostic, and scoping reviews; Human reference standard bias: Two human assessors may not represent gold standard; potential for shared interpretation biases; Training set contamination risk: Few-shot prompts developed on training set may overfit to training examples; Subset selection bias: 95 of 171 elements (55%) selected, with exclusion based on subjective difficulty assessment by one investigator; LLM model drift: Models specified as 'Preview' versions may change during study; Selection bias: Sample restricted to PubMed Central Open Access Subset indexed in September 2025 only; Selection bias: Limited to English-language systematic reviews; Selection bias: Limited to intervention reviews on human health only; Selection bias: 76 of 171 PRISMA-Check elements excluded based on principal investigator judgment of difficulty; Reference standard bias: Human consensus treated as gold standard but acknowledged as imperfect; Instruction bias: LLM and human assessors receive identical guidance and examples, potentially leading to artificial agreement unrelated to true PRISMA adherence; Operator bias: Different human assessors (DPQC assesses all 100 reviews; MYZ, DGH, BNS, PYN, MJP each assess 20 reviews); Model drift: Use of models labeled 'Preview' versions may change during study period

Limitations

  • The authors acknowledge several important limitations: "We acknowledge the inherent limitations of treating the human consensus responses as our reference standard and recognize that they do not constitute a perfect gold standard." The peer reviewer identifies additional critical limitations: "What it can actually show is which questions four LLMs answer in agreement with two trained humans using PRISMA-Check, on PubMed Central Open Access intervention reviews from one month in 2025
  • The boundaries should be stated cleanly in the discussion: English only, open access only, one indexing month, and intervention reviews of human health only
  • Paywalled reviews, non-English reviews, diagnostic accuracy reviews, prognostic reviews, scoping reviews, and non-health reviews
  • none of these are covered." The reviewer further notes: "The 76 excluded elements are a bigger issue than the protocol treats them as" and "The abstract, keywords, and conclusion should all explicitly state that this is a subset, not all of PRISMA 2020."

Open questions raised

  • The authors identify gaps addressed by their study: (1) Previous LLM validation studies assessed only PRISMA 2020 item-level checklist rather than detailed element-level; (2) Prior studies used zero-shot prompting without examples of optimal reporting; (3) Prior studies sampled only restricted health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology); (4) Future research needed to 'identify how best to consolidate the adherence assessments and to develop tools that provide constructive feedback to authors for improving their systematic reviews'
  • The protocol identifies gaps addressed in this study: (1) Prior studies assessed adherence only at the item level rather than the more detailed element level; (2) Prior studies used zero-shot prompting without examples, potentially hampering performance; (3) Prior studies restricted samples to particular health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology); (4) Prior studies did not standardize assessments using a validated tool. Future research directions mentioned include: developing methods to automate assessment of all questions in PRISMA-Check (currently only 200 of 315 questions assessed), and identifying how to consolidate adherence assessments and provide constructive feedback to authors.
  • Previous LLM studies limited to item-level rather than element-level assessment of PRISMA adherence
  • Previous studies used zero-shot prompting without exemplars or guidance
  • Previous studies restricted to particular health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology)
  • Methods needed to automate assessment of all 315 questions in full PRISMA-Check tool
Data: Authors state: "We will make LLM prompts, the clean dataset, and analytic code publicly accessible to facilitate verification of our work and future research in this area." However, specific repository URLs are not provided in this protocol document.; The protocol states: "We will make LLM prompts, the clean dataset, and analytic code publicly accessible to facilitate verification of our work and future research in this area." However, specific repository URL or access location not provided.; 200 published systematic reviews (100 training set, 100 test set) from PubMed Central Open Access Subset; 50 Cochrane reviews used for prompt development examples; Clean dataset and analytic code to be made publicly accessible post-publicationCode: Not yet specified. The protocol commits to making "analytic code publicly accessible" but does not identify a specific repository.; Raw LLM outputs to be archived with data release (specific repository not yet identified)Extracted from: pdfAgreement 49%

Explore related topics

Related papers