12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews

Lech Madeyski, Barbara Kitchenham, Martin Shepperd · Information and Software Technology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
C
Evidence
1
Citations
6.40
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.infsof.2026.108204

Methodology & findings

Study design

Mixed-methods systematic review combining: (1) extraction and analysis of metrics from 28 additional papers plus a motivating benchmark (Delgado-Chaves et al.

Main result

The study found critical gaps in current LLM-screening evaluation practices. Across analyzed papers, "only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost." Most strikingly, "in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%." The proposed Weighted Matthews Correlation Coefficient demonstrated substantial disagreement with standard metrics: "MCC and WMCC disagree on the best LLM in 55% of evaluable studies."

Research paradigm

Empirical-quantitative with methodological recommendations

Author conclusions

The authors conclude that systematic review screening evaluations require fundamental changes in practice: "SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available."

Risk of bias

Potential selection bias in the 28 papers reviewed—representativeness unclear; Reanalysis dependent on availability and quality of data from prior studies; Sensitivity analysis findings (median crossover at w ≈ 2.7) may be specific to the studies analyzed; Generalizability of benchmarks across different biomedical and software-engineering domains unclear; Selection bias in the 28 papers reviewed (methodology for paper selection not explicitly detailed); Publication bias (tendency to publish studies with favorable LLM results); Domain specificity bias (focused on biomedical and software-engineering SRs); Potential confounding from differences in LLM versions, prompting strategies, and evaluation protocols across reviewed studies; Publication bias in the reviewed literature (studies may selectively report favorable metrics); Selection bias in choice of papers reviewed; Potential heterogeneity in how metrics are reported across different studies

Limitations

  • The authors note that their recommendations, while principled, have important constraints: "Extension to full-text screening and data extraction is principled but pending empirical validation." The study is limited by the available papers reviewed and the specific domain focus on biomedical and software-engineering systematic reviews, with sensitivity analysis supporting w=10 as a conservative default but acknowledging "median crossover at w ≈ 2.7, all < 7."

Open questions raised

  • The authors identify that extension of the proposed recommendations to full-text screening and data extraction is "principled but pending empirical validation." Additionally, the review reveals that most papers do not report adequate metrics (confusion matrices, MCC) or cost-sensitive evaluation metrics.
  • Empirical validation of WMCC in full-text screening contexts
  • Empirical validation of WMCC in data extraction tasks
  • Adoption of cost-sensitive metrics in LLM-screening evaluation studies
  • Standardization of reporting practices (confusion matrices, Lost Evidence metrics) across SR-screening studies
  • Development and validation of baselines for comparison with LLM approaches
Data: not_statedCode: not_statedExtracted from: pdfAgreement 49%

Explore related topics

Related papers