12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Collaborative large language models (LLMs) are all you need for screening in systematic reviews

Mihir Parmar, Syed Arsalan Ahmed Naqvi, Kainat Warraich, Amir Saeidi, Samarth Rawal, Kunwer Sufyan Faisal et al. · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.02.07.26345640

Methodology & findings

Study design

Empirical computational study using three large language models (GPT-4 Turbo, Claude-3 Sonnet, Gemini-Pro-1.0) evaluated on five real-world systematic review datasets (11,300 articles total).

Sample

N = 11300, 10 groups

Primary method

Descriptive statistical analysis (frequencies with relative percentages for categorical variables, means with standard deviations for continuous variables). Work Saved (WS) formula: WS = (Total Samples - (TP + FP)) / Total Samples. Work Saved over Sample (WSS) formula: WSS = WS / Loss, normalizing work saved by the number of missing included articles (loss). Performance metrics: precision for exclusion and recall for inclusion. Systematic error analysis conducted by trained reviewers using pre-specified error categorization schema stratified by PICOS elements (population, intervention, control, outcomes).

Main result

The study found that collaborative LLM approaches demonstrated substantially improved performance compared to individual models. Specifically, "the proposed collaborative approaches -simulating two reviewer settings in real world -resulted in a substantial increase in overall performance, with mean precision reaching as high as 99.9% and recall as high as 99%." Additionally, "This collaborative framework also reduced potential manual screening effort by approximately 60%."

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "This study introduces a promising framework for leveraging collaborative LLMs to automate the screening phase of systematic reviews, effectively simulating the dual-reviewer model, LLM interaction and demonstrating high precision, recall, and substantial workload reduction." They further note that "The findings suggest that LLM collaboration, with structured conflict resolution, could enhance the accuracy and consistency of evidence synthesis process, particularly for living guidelines" and that the work "lays a strong foundation for future research into scalable, trustworthy, and adaptive LLMs-based systems in clinical evidence synthesis workflows."

Risk of bias

Potential data leakage: Models may have encountered training articles during pre-training, leading to overestimation of performance; Class imbalance: Dataset composition was 93.2% excluded and 6.8% included articles, which may affect evaluation metrics interpretation; Zero-shot prompt variability: Absence of in-context examples or fine-tuning may introduce variability in model responses; Model-specific biases: Comprehension errors and reasoning errors varied by model; Gemini-P showed higher incidence of hallucination and confirmation bias; Potential data leakage from pre-training: models may have encountered test articles during pre-training, leading to overestimation of performance; Class imbalance: only 6.8% of articles were labeled as included vs 93.2% excluded, affecting interpretation of evaluation metrics; Zero-shot prompt variability: absence of specific in-context examples or fine-tuning may introduce variability; Limited geographic/clinical diversity: five datasets all focused on oncology (cancer treatment reviews); Computational resource bias: computational intensity may limit scalability in resource-constrained environments; Potential data leakage: Models may have encountered training articles during pre-training phase; Selection bias: Dataset imbalance with 93.2% excluded vs 6.8% included articles; Model-specific biases: Gemini-P showed higher comprehension errors (95.7%) and contextual incoherence; Prompt engineering bias: Zero-shot approach without few-shot examples or fine-tuning; Confirmation bias and hallucination errors noted, particularly with Gemini-P

Limitations

  • The authors state several limitations: "The dependency on proprietary LLMs, such as GPT-4T, Claude-3S, and Gemini-P, might reduce the reproducibility, as these models undergo frequent updates which might affect performance consistency." Additionally, "although we used a random sample of articles to develop our prompts, it is possible that the models may have encountered these articles during their pre-training phase, potentially leading to an overestimation of their performance in a zero-shot setting." The authors also note "despite achieving high precision and recall, these models still exclude a limited set of articles that met the inclusion criteria
  • Error analysis revealed that most errors stemmed from the models' difficulties in comprehending study design as per the inclusion criteria."

Open questions raised

  • How well LLMs perform when simulating the two-reviewer model commonly used in clinical practice (previously assessed only as single-reviewer approaches)
  • Evaluation of open-source models using the proposed collaborative approach to improve reproducibility
  • Optimization of efficiency through model distillation or computationally efficient architectures for scalability in resource-constrained environments
  • Integration of few-shot prompting or fine-tuned LLMs specifically for screening articles
  • Further exploration of LLM-LLM conversational interactions for conflict resolution and decision refinement
  • Better understanding of model reasoning and deeper investigation into model performance variability
Data: Five datasets from the living interactive evidence synthesis project covering: (i) systemic treatment options in metastatic hormone-sensitive prostate cancer (mHSPC), (ii) systemic treatment options in metastatic castration-resistant prostate cancer (mCRPC), (iii) systemic treatment options in renal cell carcinoma (RCC), (iv) systemic treatment options in advanced/metastatic hepatocellular carcinoma (mHCC), and (v) toxicity of PARP inhibitors in cancer patients. The authors state that "these fully annotated datasets, along with machine-generated decision labels and explanations, have been made publicly available."; Five systematic review datasets (11,300 articles total) from the living interactive evidence synthesis project covering: (1) metastatic hormone-sensitive prostate cancer (mHSPC), (2) metastatic castration-resistant prostate cancer (mCRPC), (3) renal cell carcinoma (RCC), (4) metastatic hepatocellular carcinoma (mHCC), and (5) toxicity of PARP inhibitors in cancer patients. The authors state these datasets "have been made publicly available, providing a valuable resource for further research."; The authors state: "These fully annotated datasets, along with machine-generated decision labels and explanations, have been made publicly available, providing a valuable resource for further research and future development efforts in this space." However, specific URLs or access links are not provided in the text.Code: Not mentioned in the provided textExtracted from: pdfAgreement 60%

Explore related topics

Related papers