Peer Review Report For: Evaluation of automated assessments of systematic review adherence to the PRISMA 2020 statement: study protocol [version 1; peer review: 1 approved with reservations]
Matthew J. Page, Evan Mayo-Wilson, Minyan Zeng, David PQ Clark, Daniel G. Hamilton, Phi-Yen Nguyen et al. · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.5256/f1000research.198806.r482712
Methodology & findings
Study design
Diagnostic test accuracy analysis framework (validation study).
Sample
N = 200, 4 groups
Primary method
Python (version 3.14.3) will be used for all analyses. Methods include: (1) Diagnostic test accuracy analysis framework with consensus human response as reference standard; (2) Calculation of frequencies and percentages for systematic review characteristics; (3) Median and interquartile range for token counts, costs, and generation times; (4) Performance metrics: percentage agreement, Gwet's Agreement Coefficient (AC), sensitivity, specificity, positive predictive value, negative predictive value, F1 score; (5) Bootstrap percentile method with 5000 replications for 95% confidence intervals; (6) Clustering of observations within systematic reviews by resampling entire reviews; (7) Systematic comparison of supporting evidence extracted by LLMs and humans
Main result
This is a study protocol, not a completed empirical study. However, the authors state: "We will determine which questions in our comprehensive tool for assessing adherence to PRISMA 2020 can be accurately automated by LLMs. This knowledge will help inform which questions need the most human oversight by meta-researchers, peer reviewers and other interest holders seeking to assess adherence." The study aims to "evaluate the performance of each LLM assessment against the reference standard by calculating percentage agreement, Gwet's Agreement Coefficient, sensitivity, specificity, positive predictive value, negative predictive value, and the F1 score."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude: "Our study will determine which questions in the PRISMA-Check tool can be accurately automated by LLMs. This knowledge will help inform which questions need the most human oversight by meta-researchers, peer reviewers and other interest holders seeking to assess adherence." They further state: "We recognise that the number of questions in our adherence tool is large and that authors and editors of a systematic review might not want a report providing many Yes/No judgements. Future research could build on our planned work to identify how best to consolidate the adherence assessments and to develop tools that provide constructive feedback to authors for improving their systematic reviews."
Risk of bias
Selection bias: Systematic reviews limited to PubMed Central Open Access Subset from September 2025, excluding paywalled reviews; Language bias: Only English-language reviews included; Indexing bias: Limited to one indexing month (September 2025); Scope bias: Limited to intervention reviews on human health; excludes diagnostic accuracy, prognostic, and scoping reviews; Reference standard bias: Human consensus treated as gold standard despite acknowledged imperfections; Exclusion bias: 76 of 171 elements excluded from PRISMA-Check assessment; reviewer notes 'Difficulty is exactly where users will want LLM help, so cutting those out makes the validation easier on the models'; Prompt development bias: Prompts developed on training set; reviewer notes 'The authors should also state plainly that no test-set review will be opened, viewed, or skimmed by anyone during prompt development'; LLM selection bias: Limited to specific LLM versions available in April 2026; 'Preview' versions subject to change; Operator bias: One investigator (DPQC) screens all 200 abstracts and conducts all 100 full-text assessments in test set; Selection bias: Limited to PubMed Central Open Access Subset, excluding paywalled reviews; Language bias: English-language reviews only; Temporal bias: Systematic reviews indexed in one month (September 2025); Population bias: Intervention reviews on human health only; excludes diagnostic accuracy, prognostic, and scoping reviews; Human reference standard bias: Two human assessors may not represent gold standard; potential for shared interpretation biases; Training set contamination risk: Few-shot prompts developed on training set may overfit to training examples; Subset selection bias: 95 of 171 elements (55%) selected, with exclusion based on subjective difficulty assessment by one investigator; LLM model drift: Models specified as 'Preview' versions may change during study; Selection bias: Sample restricted to PubMed Central Open Access Subset indexed in September 2025 only; Selection bias: Limited to English-language systematic reviews; Selection bias: Limited to intervention reviews on human health only; Selection bias: 76 of 171 PRISMA-Check elements excluded based on principal investigator judgment of difficulty; Reference standard bias: Human consensus treated as gold standard but acknowledged as imperfect; Instruction bias: LLM and human assessors receive identical guidance and examples, potentially leading to artificial agreement unrelated to true PRISMA adherence; Operator bias: Different human assessors (DPQC assesses all 100 reviews; MYZ, DGH, BNS, PYN, MJP each assess 20 reviews); Model drift: Use of models labeled 'Preview' versions may change during study period
Limitations
- The authors acknowledge several important limitations: "We acknowledge the inherent limitations of treating the human consensus responses as our reference standard and recognize that they do not constitute a perfect gold standard." The peer reviewer identifies additional critical limitations: "What it can actually show is which questions four LLMs answer in agreement with two trained humans using PRISMA-Check, on PubMed Central Open Access intervention reviews from one month in 2025
- The boundaries should be stated cleanly in the discussion: English only, open access only, one indexing month, and intervention reviews of human health only
- Paywalled reviews, non-English reviews, diagnostic accuracy reviews, prognostic reviews, scoping reviews, and non-health reviews
- none of these are covered." The reviewer further notes: "The 76 excluded elements are a bigger issue than the protocol treats them as" and "The abstract, keywords, and conclusion should all explicitly state that this is a subset, not all of PRISMA 2020."
Open questions raised
- The authors identify gaps addressed by their study: (1) Previous LLM validation studies assessed only PRISMA 2020 item-level checklist rather than detailed element-level; (2) Prior studies used zero-shot prompting without examples of optimal reporting; (3) Prior studies sampled only restricted health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology); (4) Future research needed to 'identify how best to consolidate the adherence assessments and to develop tools that provide constructive feedback to authors for improving their systematic reviews'
- The protocol identifies gaps addressed in this study: (1) Prior studies assessed adherence only at the item level rather than the more detailed element level; (2) Prior studies used zero-shot prompting without examples, potentially hampering performance; (3) Prior studies restricted samples to particular health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology); (4) Prior studies did not standardize assessments using a validated tool. Future research directions mentioned include: developing methods to automate assessment of all questions in PRISMA-Check (currently only 200 of 315 questions assessed), and identifying how to consolidate adherence assessments and provide constructive feedback to authors.
- Previous LLM studies limited to item-level rather than element-level assessment of PRISMA adherence
- Previous studies used zero-shot prompting without exemplars or guidance
- Previous studies restricted to particular health fields (acupuncture, emergency medicine, rehabilitation, ophthalmology)
- Methods needed to automate assessment of all 315 questions in full PRISMA-Check tool
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations