Fully Automated Systematic Review Generation via Large Language Models: Quality Assessment and Implications for Scientific Publishing
Liam McLaughlin, Michael S. Walz, Cade Arries · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.02.18.26346559
Methodology & findings
Study design
Comparative case study with expert panel evaluation.
Sample
N = 6, 1 group
Primary method
Wilcoxon-Mann-Whitney U test (non-parametric test chosen for robustness with low sample sizes). Statistical testing performed in Python using scipy.stats module. Graphing performed in Microsoft Excel. Manual verification of citations used keyword matching via 'Ctrl + F' shortcut. No parametric tests were used despite reporting means.
Main result
The study found that when automated systematic reviews were subjected to expert evaluation, "the reviewers demonstrated a preference for the 'Semi-Automated' review (μ = 3.66), followed by the 'Fully-Automated' review (μ = 3.4), with the human review scoring the lowest (μ = 2.6)." Additionally, the authors demonstrated that "through controlled calls to the API in this manner, the citation error rate, which was found in a previous study to be as high as 70%, can be brought to lower than 5%," specifically achieving "a citation misattribution rate of 7.06% in the 'Semi-Auto' review, and an error rate of 4.13% in our 'Fully-Auto' review."
Reports effect sizes.
Research paradigm
Empirical-pragmatic (mixed methods: computational experiment with human expert evaluation)
Author conclusions
The authors conclude: "With only cost and API limits acting as a barrier, LLM API integration in code opens potential for the development of dynamic pipelines that totally automate informational systems." They recommend: "LLM use offers significant time-saving capabilities for review production, but that a human-operator is still required for validation, error-correction, and executive synthesis of the review process." They further state: "Due to the trade-off between the amount of material able to be understood by the LLM in one prompt and its error rate, we do not recommend the writing of reviews solely through API-powered 'fully-automated pipelines'." Finally, they emphasize: "we recommend the usage of AI in contained environments or through well-validated pipelines, with careful validation by a human operator. We emphasize the importance of transparency in AI-usage."
Risk of bias
Small sample size (n=6) with modest statistical power; Limited AI literacy among reviewers (3 used AI less than monthly); Potential observer bias in manual citation error verification; Selection bias in human review choice (limited to last 3 years, peer-reviewed, freely available); Reformatting of human review to match AI format may have introduced bias; Reviewer familiarity with specific lymphoma subtype may have influenced assessments; Small sample size (n=6) limits statistical power and generalizability; Reviewer selection bias: all reviewers from University of Minnesota Hematopathology Department; Lack of AI literacy among reviewers may have influenced assessments (3 of 6 used AI less than once per month); Blinding may have been compromised: human review was formatted differently (included figures) and reviewers successfully identified it as human in some cases despite it being rated lowest quality; Single evaluation topic (Classic Hodgkin Lymphoma) limits generalizability across domains; Manual verification of citation accuracy introduces verification bias; Model-specific findings (only Sonnet 3.5 and 4.0 tested) may not generalize to other LLMs; Reviewer expectations of AI quality may bias assessments (expected AI to be sloppier than human writing); Small sample size (n=6) limiting statistical power and generalizability; Limited AI literacy among reviewers (3 used AI 'less than once per month'), potentially introducing unfamiliarity bias in AI detection; Potential blinding bias from formatting changes to human review for comparison; Selection bias in human review choice (required peer-reviewed, freely-available publication); Reviewer expectation bias regarding AI-generated text quality; Manual citation accuracy verification may have missed errors
Limitations
- The authors state: "Firstly, our sample size of six was modest, and with only one review presented to them, our results did not demonstrate significant findings about the quality of AI vs human writing but instead showed trends that in the future deserve additional investigation." They also note: "We recognize that our stringent standards for creation of the 'Fully-Automated' and 'Semi-Automated' reviews, where we altered as little as possible of what the AI generated, would not accurately represent how some users would incorporate AI into their paper writing." Additionally: "we must also acknowledge that our internal analysis for errors itself may not have captured all of the potential errors that the AI made, due to the inherent barriers of manual verification of every claim." Finally: "we recognize that our evaluation of only the Claude Sonnet 3.5 and 4.0 models will not necessarily be valid for every available LLM model."
Open questions raised
- Need for improved AI literacy among medical professionals
- Evaluation of citation accuracy and error rates across different LLM models beyond Claude Sonnet 3.5 and 4.0
- Investigation of whether LLMs progressively lose understanding with large text volumes or if there is a threshold where understanding breaks down
- Long-term implications of AI-generated review flooding academic publishing spaces
- Development of standardized ethical frameworks for AI use in academic writing
- Validation of fully-automated pipelines in more contained settings and diverse use cases
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations