LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews
Lech Madeyski, Barbara Kitchenham, Martin Shepperd · Information and Software Technology · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.infsof.2026.108204
Methodology & findings
Study design
Mixed-methods systematic review combining: (1) extraction and analysis of metrics from 28 additional papers plus a motivating benchmark (Delgado-Chaves et al.
Main result
The study found critical gaps in current LLM-screening evaluation practices. Across analyzed papers, "only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost." Most strikingly, "in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%." The proposed Weighted Matthews Correlation Coefficient demonstrated substantial disagreement with standard metrics: "MCC and WMCC disagree on the best LLM in 55% of evaluable studies."
Research paradigm
Empirical-quantitative with methodological recommendations
Author conclusions
The authors conclude that systematic review screening evaluations require fundamental changes in practice: "SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available."
Risk of bias
Potential selection bias in the 28 papers reviewed—representativeness unclear; Reanalysis dependent on availability and quality of data from prior studies; Sensitivity analysis findings (median crossover at w ≈ 2.7) may be specific to the studies analyzed; Generalizability of benchmarks across different biomedical and software-engineering domains unclear; Publication bias (tendency to publish studies with favorable LLM results); Domain specificity bias (focused on biomedical and software-engineering SRs); Potential confounding from differences in LLM versions, prompting strategies, and evaluation protocols across reviewed studies; Publication bias in the reviewed literature (studies may selectively report favorable metrics); Potential heterogeneity in how metrics are reported across different studies
Limitations
- The authors note that their recommendations, while principled, have important constraints: "Extension to full-text screening and data extraction is principled but pending empirical validation." The study is limited by the available papers reviewed and the specific domain focus on biomedical and software-engineering systematic reviews, with sensitivity analysis supporting w=10 as a conservative default but acknowledging "median crossover at w ≈ 2.7, all < 7."
Open questions raised
- The authors identify that extension of the proposed recommendations to full-text screening and data extraction is "principled but pending empirical validation." Additionally, the review reveals that most papers do not report adequate metrics (confusion matrices, MCC) or cost-sensitive evaluation metrics.
- Standardization of reporting practices (confusion matrices, Lost Evidence metrics) across SR-screening studies
- Development and validation of baselines for comparison with LLM approaches
- The authors identify that "Extension to full-text screening and data extraction is principled but pending empirical validation" and recommend that "Editors and reviewers should require these elements as routine," indicating gaps in current systematic review practice and the need for standardized evaluation protocols.
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations