Supporting Literature Reviews: A Comparison Between Human and Generative Artificial Intelligence Screening for a Scoping Review.
Tami H. Wyatt, Heather Carter-Templeton, Jordan Wrigley, Martin Kang, Rosemary Kennedy, Gregory L. Alexander et al. · PubMed · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1097/cin.0000000000001587
Methodology & findings
Study design
Retrospective cross-sectional agreement design comparing title and abstract screening decisions made by ChatGPT 3.5 with decisions made by human reviewers on 3148 articles retrieved from a literature search..
Sample
N = 3148, 2 groups
Primary method
Cross-sectional agreement analysis methodology implied but specific statistical tests, software, or measures of agreement (e.g., Cohen's kappa, sensitivity, specificity) are not detailed in the abstract.
Main result
The study found that "during title and abstract screening, the human research team excluded 2661 articles (84.5%), whereas ChatGPT 3.5 excluded 1533 articles (48.7%)," indicating substantially different screening decisions between human reviewers and the AI tool. The authors note that "although AI-assisted screening may reduce time by filtering out a portion of irrelevant studies early in the process, these efficiencies must be balanced against the depth of understanding gained through review among the human team."
Reports effect sizes.
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude that "although AI-assisted screening may reduce time by filtering out a portion of irrelevant studies early in the process, these efficiencies must be balanced against the depth of understanding gained through review among the human team. Furthermore, the dialogue and consensus-building among research team members may be diminished when AI tools are used. This reduction in scholarly engagement may limit opportunities for critical appraisal, learning, and deeper comprehension of the evidence."
Risk of bias
Selection bias: Only ChatGPT 3.5 was tested; findings may not generalize to other AI tools; Technology change bias: AI models are updated frequently; results may not reflect current versions; Operator bias: Human reviewer screening decisions may vary based on expertise and interpretation; Incomplete data handling: 6 articles excluded due to incomplete data or upload errors; Selection bias in articles retrieved and screened; Potential systematic differences in exclusion criteria between human and AI decision-making; Incomplete data and upload errors (6 articles excluded); Potential selection bias: exclusion of 6 articles due to data incompleteness or upload errors could affect representativeness; No mention of blinding or independent verification of screening decisions; AI tool (ChatGPT 3.5) operationalization and prompting methodology not detailed in abstract; Potential for human reviewer bias in the comparison group
Open questions raised
- The abstract indicates that "the accuracy and impact [of AI tools] on rigor remain uncertain," suggesting a gap in understanding how AI screening affects the rigor and quality of systematic/scoping review processes.
- The authors identify that "the accuracy and impact of AI tools on rigor remain uncertain" in scoping reviews, suggesting a need for further investigation into how AI screening affects the quality and rigor of literature review processes.
- The authors identify uncertainty regarding AI tools' accuracy and impact on rigor in scoping review processes, suggesting future research is needed to determine optimal integration of AI in literature review workflows.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations