12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Validating Large Language Models for Title-Abstract Screening in Low-Prevalence Systematic Reviews: An Environmental Science Case Study

Maximilian Nawrath, Andrea Merlina, Jemmima Knight, Sam A. Welch, Mahla Rashidian, Isabel Seifert-Dähnn · Information · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
C
Evidence
1
Citations
11.26
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/info17050501

Methodology & findings

Study design

Comparative validation study using a gold-standard screening dataset of 500 articles from a published systematic review on urban greenspaces and mental health in low-and middle-income countries.

Main result

Across models, we find that LLMs can achieve high sensitivity and substantial agreement with human reviewers, but that performance varies markedly depending on the evaluation metrics used. "GPT-4.1 correctly retrieved 7 true positive studies and 431 true negatives but incorrectly included 61 false positives with one false negative error. This yielded a total accuracy of 87.6%" with sensitivity of 0.875 (95% CI: 0.529-0.978). Gemini 2.0 Flash achieved perfect sensitivity (1.0) with all 8 true positives identified but with lower specificity (0.858) and more false positives (70). DeepSeek V3 achieved the highest specificity (0.980) and overall accuracy (97.0%) but at the cost of reduced sensitivity (0.375), missing 5 of 8 relevant studies. Critical finding: "the reliance on single indicators such as overall accuracy or Cohen's κ can yield misleading conclusions in low-prevalence screening contexts, whereas prevalence-robust agreement measures and classification metrics better reflect screening priorities."

Research paradigm

Positivist/empiricist

Author conclusions

"This study provides a rigorous, multi-metric validation of contemporary LLMs for title and abstract screening in systematic reviews, using a gold-standard human reference dataset and a conservative, zero-shot prompting framework." The authors conclude that "LLMs can achieve high sensitivity and substantial agreement with human reviewers, but that performance varies markedly depending on the evaluation metrics used." Critically, they state: "the reliance on single indicators such as overall accuracy or Cohen's κ can yield misleading conclusions in low-prevalence screening contexts, whereas prevalence-robust agreement measures and classification metrics better reflect screening priorities." They further conclude that "models optimised for recall are better aligned with systematic review objectives, even when this comes at the cost of increased false positives and downstream screening effort" and that "LLM-assisted screening is not a plug-and-play solution. Performance is contingent on prompt design, metric choice, and continued human oversight. When deployed conservatively and transparently, LLMs can meaningfully reduce screening workload while preserving review validity."

Risk of bias

Training data leakage: published gold-standard dataset (Nawrath et al. 2021) may have appeared in model training corpora; Prompt optimization bias: iterative refinement using the same 500-abstract evaluation set, potentially leading to adaptive overfitting; Proprietary API limitations: models accessed via APIs without full reproducibility or independent verification; Study design criterion omission: prompt did not operationalize study design screening, potentially inflating false positives; Narrow interpretation bias: recurrent patterns of overly restrictive interpretations of mental health and population criteria by some models; Leniency instruction bias: explicit lenient decision rule may have differentially affected false positive rates across models; Limited stability testing: single-run evaluation without repeated API calls to quantify backend nondeterminism; Data leakage: Prompt refinement conducted on the same 500-abstract evaluation corpus; Training data contamination: Openly published dataset may have appeared in model training corpora; Prevalence bias: Severe class imbalance (8 inclusions vs. 492 exclusions) affects performance metrics; Prompt overfitting: Iterative refinement may have optimized performance to this specific dataset; Model drift: Performance characteristics may change with API backend updates; Selection bias in leniency rule: Explicit instruction to be lenient may systematically bias toward false positives; Criterion operationalization bias: Study design not included as screening criterion, potentially inflating false positives; Training data leakage: openly available gold-standard dataset (Nawrath et al. 2021) may have appeared in LLM training corpora; Adaptive overfitting: prompt refinement drew on outputs from the same 500-abstract corpus used for evaluation; Selection bias: only title-abstract screening evaluated, not full-text screening; Low prevalence effects: only 8 included studies out of 500, making sensitivity estimates unstable; API non-determinism: backend variation in proprietary LLM APIs despite temperature=0.0 setting; Prompt optimization bias: performance reported as conditional on optimized prompt, not intrinsic model property

Limitations

  • "Because we refined the prompt iteratively by inspecting the outputs from the evaluated corpus of studies, results may be modestly optimistic due to overfitting to a reused test set." Additional limitations include: "the gold-standard dataset was published openly in 2021, meaning records may have appeared in the training corpora of the evaluated models
  • We cannot fully rule this out, as training data for proprietary models are not publicly disclosed." Furthermore, "all five models were accessed via proprietary APIs, limiting reproducibility and independent verification, which is a recognised challenge in LLM evaluation for evidence synthesis." The study did not conduct "repeated-run stability testing such as multiple identical API calls per model, to quantify the run-to-run variability under otherwise fixed settings." Additionally, "sensitivity estimates and odds-ratios reported in this study should be interpreted with caution, given that small changes in classification would materially affect these results" due to the small number of inclusions (n=8).

Open questions raised

  • Lack of generalizability: Most validation studies focus on biomedical and environmental sciences, leaving gaps in social sciences, engineering, and interdisciplinary research
  • Limited multilingual evaluation: Expanding to multilingual and non-English corpora essential, as language bias may affect screening accuracy and equity
  • Reproducibility gaps: Need for explicit testing of stability across repeated runs, model versions, access modalities, and time; sensitivity to prompt variation underexplored
  • Methodological standardization: Clearer reporting standards needed specifying model version, access modality, prompt strategy, and evaluation metrics
  • Full-text screening validation: Validation has focused on title-abstract screening; full-text screening where decisions are more complex remains underexplored
  • Prompt variation exploration: Systematic exploration of prompt formats, chain-of-thought strategies, and domain adaptation limited
Data: Full dataset publicly available at https://github.com/NIVANorge/ai-literature-review-public/blob/main/Full%20dataset.xlsx (accessed on 1 May 2025), including all model decisions and reasoning; Gold-standard screening dataset derived from Nawrath et al. (2021) - 500-article randomly selected subset from 1801 potentially eligible articles on urban greenspaces and mental health in low-and middle-income countries; Gold-standard screening dataset; Gold-standard dataset from Nawrath et al. (2021) on urban greenspaces and mental health outcomes in LMICs; Full evaluation dataset: https://github.com/NIVANorge/ai-literature-review-public/blob/main/Full%20dataset.xlsx (accessed on 1 May 2025)Code: Python code (version 3.11.9) used to generate results is openly available in accompanying repository: https://github.com/NIVANorge/ai-literature-review-public/ (referenced as [28] in paper); https://github.com/NIVANorge/ai-literaturereview-public; Python code repository: https://github.com/NIVANorge/ai-literature-review-public/ (Python version 3.11.9, code is openly available in the accompanying repository)Extracted from: pdfAgreement 36%

Explore related topics

Related papers