Automating abstract screening in research synthesis using large language models: A tutorial and proof-of-concept study
Mirka Henninger, Jan Radek, Jean-Paul Snijder, Martin Pauly · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.31234/osf.io/pjmcs_v4
Methodology & findings
Study design
Proof-of-concept study with tutorial development.
Main result
The authors demonstrate that "LLM-assisted screening can substantially reduce the time and cost of review preparation while maintaining accuracy comparable to human raters," as illustrated through their proof-of-concept study using human-rated abstracts from a recent meta-analysis for validation.
Research paradigm
pragmatist/applied computational
Author conclusions
The authors conclude that they "provide a step-by-step tutorial in which [they] address how to select an LLM, access LLMs from within R, develop and refine suitable prompts, define structured output formats, and whether and how model hyperparameters should be set," with the goal of supporting psychological researchers in adopting LLM-based workflows for research synthesis.
Risk of bias
Single proof-of-concept validation set; no discussion of inter-rater reliability comparison methodology; potential selection bias in choice of meta-analysis abstracts for validation; no discussion of potential LLM hallucination or consistency issues across multiple screening runs.
Limitations
- The authors emphasize that "this work represents an initial step, and that continued refinement and validation are essential as LLM technologies and their applications continue to evolve rapidly," indicating that the current study has preliminary status and requires further validation as technology advances.
Open questions raised
- The authors identify that "there is still little guidance and practical insight into how such automated workflows can be implemented in research practice" and emphasize that continued refinement and validation of LLM applications in abstract screening are essential as the technology evolves.
- Need for continued refinement and validation of LLM-based abstract screening as technologies evolve; need for practical guidance and implementation examples in research practice; need for uncertainty quantification in LLM model outputs.
- The authors identify the gap of "little guidance and practical insight into how such automated workflows can be implemented in research practice" and note that continued refinement and validation are essential as LLM technologies evolve.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations