12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviews

Lech Madeyski, Barbara Kitchenham, Martin Shepperd · Information and Software Technology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
C
Evidence
1
Citations
6.40
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.infsof.2026.108204

Methodology & findings

Study design

Mixed-methods systematic review combining: (1) extraction and analysis of metrics from 28 additional papers plus a motivating benchmark (Delgado-Chaves et al.

Main result

The study found critical gaps in current LLM-screening evaluation practices. Across analyzed papers, "only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost." Most strikingly, "in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%." The proposed Weighted Matthews Correlation Coefficient demonstrated substantial disagreement with standard metrics: "MCC and WMCC disagree on the best LLM in 55% of evaluable studies."

Research paradigm

Empirical-quantitative with methodological recommendations

Author conclusions

The authors conclude that systematic review screening evaluations require fundamental changes in practice: "SR-screening evaluations should prioritize Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available."

Risk of bias

Potential selection bias in the 28 papers reviewed—representativeness unclear; Reanalysis dependent on availability and quality of data from prior studies; Sensitivity analysis findings (median crossover at w ≈ 2.7) may be specific to the studies analyzed; Generalizability of benchmarks across different biomedical and software-engineering domains unclear; Publication bias (tendency to publish studies with favorable LLM results); Domain specificity bias (focused on biomedical and software-engineering SRs); Potential confounding from differences in LLM versions, prompting strategies, and evaluation protocols across reviewed studies; Publication bias in the reviewed literature (studies may selectively report favorable metrics); Potential heterogeneity in how metrics are reported across different studies

Limitations

  • The authors note that their recommendations, while principled, have important constraints: "Extension to full-text screening and data extraction is principled but pending empirical validation." The study is limited by the available papers reviewed and the specific domain focus on biomedical and software-engineering systematic reviews, with sensitivity analysis supporting w=10 as a conservative default but acknowledging "median crossover at w ≈ 2.7, all < 7."

Open questions raised

  • The authors identify that extension of the proposed recommendations to full-text screening and data extraction is "principled but pending empirical validation." Additionally, the review reveals that most papers do not report adequate metrics (confusion matrices, MCC) or cost-sensitive evaluation metrics.
  • Standardization of reporting practices (confusion matrices, Lost Evidence metrics) across SR-screening studies
  • Development and validation of baselines for comparison with LLM approaches
  • The authors identify that "Extension to full-text screening and data extraction is principled but pending empirical validation" and recommend that "Editors and reviewers should require these elements as routine," indicating gaps in current systematic review practice and the need for standardized evaluation protocols.
Data: not_statedCode: not_statedExtracted from: pdf

Explore related topics

Related papers