12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Erika Yahata · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Controlled empirical evaluation comparing 12 LLMs from 4 providers (OpenAI, Google DeepMind, Anthropic, Llama AI) and 4 classical machine learning classifiers (Multinomial Naive Bayes, Logistic Regression, Random Forest, SVM) on study screening tasks across 2 real Systematic Literature Reviews (SLRs).

Sample

N = 518, 9 groups

Primary method

Bootstrap confidence intervals (Tibshirani and Efron [1993]) - 95% CI stratified across 5 iterations; Gwet's AC2 coefficient with quadratic weighting for measuring agreement across iterations (chosen over Cohen's Kappa due to robustness in class-imbalanced settings); DerSimonian-Laird random-effects model for meta-analysis of feature composition effects; TF-IDF vectorization for classical ML preprocessing; Stratified 4-fold cross-validation for classical model training; Smallest Effect Size of Interest (SESOI) calculation set to ±2.0 percentage points of accuracy

Main result

The study found substantial heterogeneity across LLMs in screening performance, with "Even with temperature set to zero, several LLMs produced non-deterministic screening decisions, indicating that stability should be treated as an empirical, context-dependent requirement rather than an operational assumption." Additionally, "abstracts emerged as the decisive metadata element for LLM-based screening: removing abstracts consistently degraded performance, whereas adding titles and keywords to abstracts provided no robust practical gains." Furthermore, "LLMs and classical methods exhibited broadly comparable performance under the same protocol, with overlapping bootstrap confidence intervals."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-positivist (controlled comparative evaluation with quantitative metrics)

Author conclusions

"Overall, the study clarifies practical limits and conditions for using LLMs in screening, showing that model scale and version do not guarantee reliability, that residual variability can persist under controlled settings, and that abstract availability is critical." The authors further conclude that "method selection should be guided primarily by operational and scientific governance criteria, including cost, transparency, reproducibility, and metadata availability, rather than by aggregated performance metrics alone." They emphasize that "LLMs and classical methods exhibited broadly comparable performance under the same protocol, with overlapping bootstrap confidence intervals."

Risk of bias

Temporal dependence: API-based model evaluations (May-October 2025) subject to model updates affecting reproducibility; Limited external validity: Only 2 SLRs from educational domains; results may not generalize to other SE domains or multilingual corpora; Class imbalance effects: SLR1 had 61.1% excluded vs 38.9% included; SLR2 had 67.6% excluded vs 32.4% included, potentially inflating accuracy metrics; Reference label reliability: Authors acknowledge screening errors may partly reflect human ambiguity rather than exclusively model failures; Variability underestimation: Only 5 iterations per model may underestimate variance in unstable settings; Selection bias in dataset preparation: 8% of records lost in SLR1 (8/134) and 12.5% in SLR2 (56/448) due to missing metadata; Selection bias in dataset construction: missing abstracts and keywords reduced samples (94% retrieval for SLR1, 87.5% for SLR2); Domain specificity: results limited to educational domains (HCI-AI, Educational Gamification); Temporal dependence: API-accessed models (May-October 2025) subject to version updates affecting reproducibility; Class imbalance sensitivity: differential performance based on inclusion/exclusion proportions in datasets; Screening error attribution ambiguity: human annotation errors not distinguished from model failures; Limited iteration count: 5 iterations may underestimate variance in unstable settings; Temporal dependence: API-based models evaluated during specific period (May-October 2025) may differ when updated; Domain-specific selection bias: Only 2 SLRs in educational domains (HCI-AI, Educational Gamification); Language bias: Only English-language publications included; Limited iterations: 5 iterations per model may underestimate variability in unstable settings; Reference label ambiguity: Screening errors may reflect human ambiguity in original reference labels rather than model failures; Class imbalance: Low inclusion rates typical in SLR screening may inflate accuracy metrics; Training set bias: Classical models trained on reduced set of 50 studies, not full datasets

Limitations

  • The authors state: "External validity: Results are limited to 2 SLRs in educational domains (HCI-AI, Educational Gamification) published in English
  • Patterns may not generalize to other Software Engineering domains or multilingual corpora." Additionally, "Internal validity: Models evaluated via APIs (May-October 2025) introduce temporal dependence
  • subsequent updates may affect reproducibility." Furthermore, "Construct validity: Variability estimated from 5 iterations per model suffices to evidence non-determinism but may underestimate variance in unstable settings." And "Conclusion validity: LLM sample excluded fine-tuned or specialized models
  • no cost-benefit analysis was performed, limiting prescriptive recommendations in resource-constrained contexts."

Open questions raised

  • Future work should: (1) expand to additional Software Engineering domains and multilingual corpora; (2) strengthen reproducibility analyses in API-based settings by tracking versions and execution windows; (3) increase the number of repeated runs to better characterize variability; (4) incorporate cost-benefit analyses at scale alongside assessments of reference-label reliability; (5) investigate fine-tuned or specialized models; (6) perform comprehensive cost-benefit analyses in resource-constrained contexts.
  • Expansion to additional Software Engineering domains beyond HCI-AI and Educational Gamification
  • Extension to multilingual corpora (current study limited to English)
  • Strengthened reproducibility analyses in API-based settings by tracking versions and execution windows
  • Increase in the number of repeated runs to better characterize variability (beyond 5 iterations)
  • Incorporation of cost-benefit analyses at scale alongside assessments of reference-label reliability
Data: SLR1: 126 complete records from tertiary study on HCI-AI convergence (Travassos et al. 2017); SLR2: 392 complete records from systematic mapping study on user classification strategies in game-based and gamified learning environments (Pessoa et al. 2024). Original datasets referenced from Felizardo et al. [2024].; SLR1: 126 studies from tertiary study on HCI-AI convergence (Travassos et al. [2017]); SLR2: 392 studies from systematic mapping study on user classification strategies in game-based and gamified learning environments (Pessoa et al. [2024]); SLR1: HCI-AI convergence dataset from Travassos et al. [2017] - 126 studies with consistent metadata (94% retrieval rate from original 134); SLR2: User classification in game-based learning dataset from Pessoa et al. [2024] - 392 studies with consistent metadata (87.5% retrieval rate from original 448)Extracted from: pdfAgreement 55%

Explore related topics

Related papers