Collaborative large language models (LLMs) are all you need for screening in systematic reviews
Mihir Parmar, Syed Arsalan Ahmed Naqvi, Kainat Warraich, Amir Saeidi, Samarth Rawal, Kunwer Sufyan Faisal et al. · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.02.07.26345640
Methodology & findings
Study design
Empirical computational study using three large language models (GPT-4 Turbo, Claude-3 Sonnet, Gemini-Pro-1.0) evaluated on five real-world systematic review datasets (11,300 articles total).
Sample
N = 11300, 10 groups
Primary method
Descriptive statistical analysis (frequencies with relative percentages for categorical variables, means with standard deviations for continuous variables). Work Saved (WS) formula: WS = (Total Samples - (TP + FP)) / Total Samples. Work Saved over Sample (WSS) formula: WSS = WS / Loss, normalizing work saved by the number of missing included articles (loss). Performance metrics: precision for exclusion and recall for inclusion. Systematic error analysis conducted by trained reviewers using pre-specified error categorization schema stratified by PICOS elements (population, intervention, control, outcomes).
Main result
The study found that collaborative LLM approaches demonstrated substantially improved performance compared to individual models. Specifically, "the proposed collaborative approaches -simulating two reviewer settings in real world -resulted in a substantial increase in overall performance, with mean precision reaching as high as 99.9% and recall as high as 99%." Additionally, "This collaborative framework also reduced potential manual screening effort by approximately 60%."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "This study introduces a promising framework for leveraging collaborative LLMs to automate the screening phase of systematic reviews, effectively simulating the dual-reviewer model, LLM interaction and demonstrating high precision, recall, and substantial workload reduction." They further note that "The findings suggest that LLM collaboration, with structured conflict resolution, could enhance the accuracy and consistency of evidence synthesis process, particularly for living guidelines" and that the work "lays a strong foundation for future research into scalable, trustworthy, and adaptive LLMs-based systems in clinical evidence synthesis workflows."
Risk of bias
Potential data leakage: Models may have encountered training articles during pre-training, leading to overestimation of performance; Class imbalance: Dataset composition was 93.2% excluded and 6.8% included articles, which may affect evaluation metrics interpretation; Zero-shot prompt variability: Absence of in-context examples or fine-tuning may introduce variability in model responses; Model-specific biases: Comprehension errors and reasoning errors varied by model; Gemini-P showed higher incidence of hallucination and confirmation bias; Potential data leakage from pre-training: models may have encountered test articles during pre-training, leading to overestimation of performance; Class imbalance: only 6.8% of articles were labeled as included vs 93.2% excluded, affecting interpretation of evaluation metrics; Zero-shot prompt variability: absence of specific in-context examples or fine-tuning may introduce variability; Limited geographic/clinical diversity: five datasets all focused on oncology (cancer treatment reviews); Computational resource bias: computational intensity may limit scalability in resource-constrained environments; Potential data leakage: Models may have encountered training articles during pre-training phase; Selection bias: Dataset imbalance with 93.2% excluded vs 6.8% included articles; Model-specific biases: Gemini-P showed higher comprehension errors (95.7%) and contextual incoherence; Prompt engineering bias: Zero-shot approach without few-shot examples or fine-tuning; Confirmation bias and hallucination errors noted, particularly with Gemini-P
Limitations
- The authors state several limitations: "The dependency on proprietary LLMs, such as GPT-4T, Claude-3S, and Gemini-P, might reduce the reproducibility, as these models undergo frequent updates which might affect performance consistency." Additionally, "although we used a random sample of articles to develop our prompts, it is possible that the models may have encountered these articles during their pre-training phase, potentially leading to an overestimation of their performance in a zero-shot setting." The authors also note "despite achieving high precision and recall, these models still exclude a limited set of articles that met the inclusion criteria
- Error analysis revealed that most errors stemmed from the models' difficulties in comprehending study design as per the inclusion criteria."
Open questions raised
- How well LLMs perform when simulating the two-reviewer model commonly used in clinical practice (previously assessed only as single-reviewer approaches)
- Evaluation of open-source models using the proposed collaborative approach to improve reproducibility
- Optimization of efficiency through model distillation or computationally efficient architectures for scalability in resource-constrained environments
- Integration of few-shot prompting or fine-tuned LLMs specifically for screening articles
- Further exploration of LLM-LLM conversational interactions for conflict resolution and decision refinement
- Better understanding of model reasoning and deeper investigation into model performance variability
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations