Using full agreement across multiple large language models for title-and-abstract screening in systematic reviews: a proof-of-concept
Frederic Hilkenmeier, Merle Stoltenberg, Christian Stierle · Systematic Reviews · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s13643-026-03228-4
Methodology & findings
Study design
Proof-of-concept validation study combining multiple large language models (LLMs) for automated title-and-abstract screening in systematic reviews.
Sample
N = 1260, 9 groups
Primary method
Krippendorff's alpha (reliability measure accounting for chance agreement); F1-score (harmonic mean of precision and recall); Accuracy; Precision; Recall/Sensitivity; Specificity; one-tailed paired t-tests (comparing metrics between majority-vote and full agreement approaches); Wilcoxon test (non-parametric alternative); sign test (conservative non-parametric test). Software used: Google Sheets with Chrome extensions for API integration to LLMs; calculations performed within spreadsheet environment.
Main result
The study found that "limiting automated title-and-abstract screening decisions to cases of full agreement across three independent LLMs was associated with consistently better performance than previously used automated approaches." Specifically, the full agreement approach achieved "Krippendorff 's alpha = 0.94 and F1-score = 0.95" on the case study calibration sample, with "the full agreement classification deviated from the consensus-based screening decisions in only 5 of 402 cases." Across six datasets examined, "One-tailed paired t-tests comparing the metrics in Tables 1 and 2 yield p-values < 0.05, confirming a statistically significant improvement."
Reports effect sizes.
Research paradigm
Empirical-positivist (quantitative validation across datasets)
Author conclusions
"In conclusion, this study suggests that limiting automated title-and-abstract screening decisions to cases of full agreement across three independent LLMs may offer a promising and conservative approach to LLM-assisted screening in systematic reviews. Across the six datasets examined, the approach was associated with better performance than single-LLM screening and majority-vote aggregation, albeit for a reduced subset of automatically classified records. The proposed workflow is therefore best understood not as a replacement for human reviewers, but as a proof-of-concept for reducing screening burden while preserving human oversight at critical stages."
Risk of bias
Selection bias in dataset selection (only six datasets examined, four from prior studies); Limited generalizability across review topics, disciplines, and non-English records; Sensitivity to prompt design—22 iterations required in case study suggests potential for researcher bias in prompt refinement; Content filtering by Gemini Pro 1.0 prevented classification of 15 records with sensitive terms, introducing potential missing data bias; Dependence on proprietary LLM platforms subject to version changes and content-moderation behavior; Calibration sample from authors' own case study may not be representative of other review contexts; Calibration sample bias: performance on 10-20% calibration subset may not reflect performance on full dataset; Single case study context: measles review topic specificity limits transferability; Potential selection bias in choosing the six datasets for analysis; Calibration sample bias: human reviewer decisions may not represent true population of relevant studies; Prompt engineering bias: iterative refinement of prompts on calibration sample could lead to overfitting; Single-reviewer validation bias: supplementary analysis uses single reviewer as reference, not consensus; Model architecture bias: choice of three specific LLM families may not generalize to other models
Limitations
- The authors acknowledge several limitations: "First, the performance of the proposed workflow depends substantially on prompt design and prompt calibration
- In our case example, 22 prompt iterations were required before the predefined performance thresholds were met." Second, "the efficiency gains of the proposed workflow are likely to vary across review contexts
- Topics with more ambiguous eligibility criteria, more heterogeneous abstracts, or more borderline cases may yield lower full-agreement rates and therefore smaller reductions in manual screening effort." Third, "the broader generalizability of the present findings remains limited
- The study examined six datasets, including reanalyzed datasets from prior work and two additional datasets from our own case study, but this evidence base remains relatively small." Additionally, "because the workflow relies on proprietary general-purpose LLMs, future work should examine how version changes, content-moderation behavior, and other platform-specific constraints may affect reproducibility across review settings and over time."
Open questions raised
- Alternative prompting strategies (few-shot prompting) should be examined to reduce setup effort and improve robustness across review contexts
- Evaluation across additional review topics, disciplines, implementation settings, and non-English records is needed
- Future work should examine how version changes, content-moderation behavior, and other platform-specific constraints affect reproducibility across review settings and over time
- Research on generalizability across topics with more ambiguous eligibility criteria, heterogeneous abstracts, and borderline cases
- Investigation of how ambient criteria ambiguity and abstract heterogeneity affect full-agreement rates
- Further validation studies to inform best-practice guidance as generative AI tools develop rapidly
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations