Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs
Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Erika Yahata · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Controlled empirical evaluation comparing 12 LLMs from 4 providers (OpenAI, Google DeepMind, Anthropic, Llama AI) and 4 classical machine learning classifiers (Multinomial Naive Bayes, Logistic Regression, Random Forest, SVM) on study screening tasks across 2 real Systematic Literature Reviews (SLRs).
Sample
N = 518, 9 groups
Primary method
Bootstrap confidence intervals (Tibshirani and Efron [1993]) - 95% CI stratified across 5 iterations; Gwet's AC2 coefficient with quadratic weighting for measuring agreement across iterations (chosen over Cohen's Kappa due to robustness in class-imbalanced settings); DerSimonian-Laird random-effects model for meta-analysis of feature composition effects; TF-IDF vectorization for classical ML preprocessing; Stratified 4-fold cross-validation for classical model training; Smallest Effect Size of Interest (SESOI) calculation set to ±2.0 percentage points of accuracy
Main result
The study found substantial heterogeneity across LLMs in screening performance, with "Even with temperature set to zero, several LLMs produced non-deterministic screening decisions, indicating that stability should be treated as an empirical, context-dependent requirement rather than an operational assumption." Additionally, "abstracts emerged as the decisive metadata element for LLM-based screening: removing abstracts consistently degraded performance, whereas adding titles and keywords to abstracts provided no robust practical gains." Furthermore, "LLMs and classical methods exhibited broadly comparable performance under the same protocol, with overlapping bootstrap confidence intervals."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-positivist (controlled comparative evaluation with quantitative metrics)
Author conclusions
"Overall, the study clarifies practical limits and conditions for using LLMs in screening, showing that model scale and version do not guarantee reliability, that residual variability can persist under controlled settings, and that abstract availability is critical." The authors further conclude that "method selection should be guided primarily by operational and scientific governance criteria, including cost, transparency, reproducibility, and metadata availability, rather than by aggregated performance metrics alone." They emphasize that "LLMs and classical methods exhibited broadly comparable performance under the same protocol, with overlapping bootstrap confidence intervals."
Risk of bias
Temporal dependence: API-based model evaluations (May-October 2025) subject to model updates affecting reproducibility; Limited external validity: Only 2 SLRs from educational domains; results may not generalize to other SE domains or multilingual corpora; Class imbalance effects: SLR1 had 61.1% excluded vs 38.9% included; SLR2 had 67.6% excluded vs 32.4% included, potentially inflating accuracy metrics; Reference label reliability: Authors acknowledge screening errors may partly reflect human ambiguity rather than exclusively model failures; Variability underestimation: Only 5 iterations per model may underestimate variance in unstable settings; Selection bias in dataset preparation: 8% of records lost in SLR1 (8/134) and 12.5% in SLR2 (56/448) due to missing metadata; Selection bias in dataset construction: missing abstracts and keywords reduced samples (94% retrieval for SLR1, 87.5% for SLR2); Domain specificity: results limited to educational domains (HCI-AI, Educational Gamification); Temporal dependence: API-accessed models (May-October 2025) subject to version updates affecting reproducibility; Class imbalance sensitivity: differential performance based on inclusion/exclusion proportions in datasets; Screening error attribution ambiguity: human annotation errors not distinguished from model failures; Limited iteration count: 5 iterations may underestimate variance in unstable settings; Temporal dependence: API-based models evaluated during specific period (May-October 2025) may differ when updated; Domain-specific selection bias: Only 2 SLRs in educational domains (HCI-AI, Educational Gamification); Language bias: Only English-language publications included; Limited iterations: 5 iterations per model may underestimate variability in unstable settings; Reference label ambiguity: Screening errors may reflect human ambiguity in original reference labels rather than model failures; Class imbalance: Low inclusion rates typical in SLR screening may inflate accuracy metrics; Training set bias: Classical models trained on reduced set of 50 studies, not full datasets
Limitations
- The authors state: "External validity: Results are limited to 2 SLRs in educational domains (HCI-AI, Educational Gamification) published in English
- Patterns may not generalize to other Software Engineering domains or multilingual corpora." Additionally, "Internal validity: Models evaluated via APIs (May-October 2025) introduce temporal dependence
- subsequent updates may affect reproducibility." Furthermore, "Construct validity: Variability estimated from 5 iterations per model suffices to evidence non-determinism but may underestimate variance in unstable settings." And "Conclusion validity: LLM sample excluded fine-tuned or specialized models
- no cost-benefit analysis was performed, limiting prescriptive recommendations in resource-constrained contexts."
Open questions raised
- Future work should: (1) expand to additional Software Engineering domains and multilingual corpora; (2) strengthen reproducibility analyses in API-based settings by tracking versions and execution windows; (3) increase the number of repeated runs to better characterize variability; (4) incorporate cost-benefit analyses at scale alongside assessments of reference-label reliability; (5) investigate fine-tuned or specialized models; (6) perform comprehensive cost-benefit analyses in resource-constrained contexts.
- Expansion to additional Software Engineering domains beyond HCI-AI and Educational Gamification
- Extension to multilingual corpora (current study limited to English)
- Strengthened reproducibility analyses in API-based settings by tracking versions and execution windows
- Increase in the number of repeated runs to better characterize variability (beyond 5 iterations)
- Incorporation of cost-benefit analyses at scale alongside assessments of reference-label reliability
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations