Using full agreement across multiple large language models for title-and-abstract screening in systematic reviews: a proof-of-concept
Frederic Hilkenmeier, Merle Stoltenberg, Christian Stierle · Systematic Reviews · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s13643-026-03228-4
Methodology & findings
Study design
Proof-of-concept validation study combining multiple large language models (LLMs) for automated title-and-abstract screening in systematic reviews.
Sample
N = 1260, 14 groups
Primary method
Krippendorff's alpha (reliability measure accounting for chance agreement); F1-score (harmonic mean of precision and recall); Accuracy; Precision; Recall/Sensitivity; Specificity; one-tailed paired t-tests (comparing metrics between majority-vote and full agreement approaches); Wilcoxon test (non-parametric alternative); sign test (conservative non-parametric test). Software used: Google Sheets with Chrome extensions for API integration to LLMs; calculations performed within spreadsheet environment.
Main result
The study found that "limiting automated title-and-abstract screening decisions to cases of full agreement across three independent LLMs was associated with consistently better performance than previously used automated approaches." Specifically, the full agreement approach achieved "Krippendorff 's alpha = 0.94 and F1-score = 0.95" on the case study calibration sample, with "the full agreement classification deviated from the consensus-based screening decisions in only 5 of 402 cases." Across six datasets examined, "One-tailed paired t-tests comparing the metrics in Tables 1 and 2 yield p-values < 0.05, confirming a statistically significant improvement."
Reports effect sizes.
Research paradigm
Empirical-positivist (quantitative validation across datasets)
Author conclusions
"In conclusion, this study suggests that limiting automated title-and-abstract screening decisions to cases of full agreement across three independent LLMs may offer a promising and conservative approach to LLM-assisted screening in systematic reviews. Across the six datasets examined, the approach was associated with better performance than single-LLM screening and majority-vote aggregation, albeit for a reduced subset of automatically classified records. The proposed workflow is therefore best understood not as a replacement for human reviewers, but as a proof-of-concept for reducing screening burden while preserving human oversight at critical stages."
Risk of bias
Selection bias in dataset selection (only six datasets examined, four from prior studies); Limited generalizability across review topics, disciplines, and non-English records; Sensitivity to prompt design—22 iterations required in case study suggests potential for researcher bias in prompt refinement; Content filtering by Gemini Pro 1.0 prevented classification of 15 records with sensitive terms, introducing potential missing data bias; Dependence on proprietary LLM platforms subject to version changes and content-moderation behavior; Calibration sample from authors' own case study may not be representative of other review contexts; Limited generalizability: only 6 datasets examined; Potential prompt engineering bias: 22 iterations required in case study may not reflect typical effort; LLM version/platform dependencies: proprietary models subject to changes beyond researchers' control; Calibration sample bias: performance on 10-20% calibration subset may not reflect performance on full dataset; Content filtering effects: Gemini Pro 1.0 refused to classify 15 records with sensitive terms, introducing systematic exclusion bias; Single case study context: measles review topic specificity limits transferability; Potential selection bias in choosing the six datasets for analysis; Calibration sample bias: human reviewer decisions may not represent true population of relevant studies; Content filtering bias: Gemini Pro 1.0 refused to classify 15 records due to sensitive terms, potentially introducing systematic exclusion; Prompt engineering bias: iterative refinement of prompts on calibration sample could lead to overfitting; Single-reviewer validation bias: supplementary analysis uses single reviewer as reference, not consensus; Model architecture bias: choice of three specific LLM families may not generalize to other models
Limitations
- The authors acknowledge several limitations: "First, the performance of the proposed workflow depends substantially on prompt design and prompt calibration
- In our case example, 22 prompt iterations were required before the predefined performance thresholds were met." Second, "the efficiency gains of the proposed workflow are likely to vary across review contexts
- Topics with more ambiguous eligibility criteria, more heterogeneous abstracts, or more borderline cases may yield lower full-agreement rates and therefore smaller reductions in manual screening effort." Third, "the broader generalizability of the present findings remains limited
- The study examined six datasets, including reanalyzed datasets from prior work and two additional datasets from our own case study, but this evidence base remains relatively small." Additionally, "because the workflow relies on proprietary general-purpose LLMs, future work should examine how version changes, content-moderation behavior, and other platform-specific constraints may affect reproducibility across review settings and over time."
Open questions raised
- Alternative prompting strategies (few-shot prompting) should be examined to reduce setup effort and improve robustness across review contexts
- Evaluation across additional review topics, disciplines, implementation settings, and non-English records is needed
- Future work should examine how version changes, content-moderation behavior, and other platform-specific constraints affect reproducibility across review settings and over time
- Research on generalizability across topics with more ambiguous eligibility criteria, heterogeneous abstracts, and borderline cases
- Need for evaluation across additional review topics, disciplines, implementation settings, and non-English records
- Impact of version changes, content-moderation behavior, and platform-specific constraints on reproducibility
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations