LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews
Lech Madeyski, Barbara Kitchenham, Martin Shepperd · Information and Software Technology · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.infsof.2026.108204
Methodology & findings
Study design
Mixed-methods systematic review combining: (1) extraction and analysis of metrics from 28 additional papers plus a motivating benchmark (Delgado-Chaves et al.
Main result
The study found critical gaps in current LLM-screening evaluation practices. Across analyzed papers, "only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost." Most strikingly, "in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%." The proposed Weighted Matthews Correlation Coefficient demonstrated substantial disagreement with standard metrics: "MCC and WMCC disagree on the best LLM in 55% of evaluable studies."
Research paradigm
Empirical-quantitative with methodological recommendations
Author conclusions
The authors conclude that systematic review screening evaluations require fundamental changes in practice: "SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available."
Risk of bias
Potential selection bias in the 28 papers reviewed—representativeness unclear; Reanalysis dependent on availability and quality of data from prior studies; Sensitivity analysis findings (median crossover at w ≈ 2.7) may be specific to the studies analyzed; Generalizability of benchmarks across different biomedical and software-engineering domains unclear; Selection bias in the 28 papers reviewed (methodology for paper selection not explicitly detailed); Publication bias (tendency to publish studies with favorable LLM results); Domain specificity bias (focused on biomedical and software-engineering SRs); Potential confounding from differences in LLM versions, prompting strategies, and evaluation protocols across reviewed studies; Publication bias in the reviewed literature (studies may selectively report favorable metrics); Selection bias in choice of papers reviewed; Potential heterogeneity in how metrics are reported across different studies
Limitations
- The authors note that their recommendations, while principled, have important constraints: "Extension to full-text screening and data extraction is principled but pending empirical validation." The study is limited by the available papers reviewed and the specific domain focus on biomedical and software-engineering systematic reviews, with sensitivity analysis supporting w=10 as a conservative default but acknowledging "median crossover at w ≈ 2.7, all < 7."
Open questions raised
- The authors identify that extension of the proposed recommendations to full-text screening and data extraction is "principled but pending empirical validation." Additionally, the review reveals that most papers do not report adequate metrics (confusion matrices, MCC) or cost-sensitive evaluation metrics.
- Empirical validation of WMCC in full-text screening contexts
- Empirical validation of WMCC in data extraction tasks
- Adoption of cost-sensitive metrics in LLM-screening evaluation studies
- Standardization of reporting practices (confusion matrices, Lost Evidence metrics) across SR-screening studies
- Development and validation of baselines for comparison with LLM approaches
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations