12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations

Mika Mäntylä, Patrícia Matsubara, Katia Romero Felizardo, Miikka Kuutila, Marco Gerosa, Savio de Sousa Sampaio et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Qualitative cross-study analysis of disagreements between LLMs (gemini-2.5-flash and openai/gpt-4.1-mini) and human raters across six software engineering systematic reviews.

Sample

N = 1044, 6 groups

Primary method

Cohen's Kappa for inter-rater agreement between human consensus and LLM decisions. Qualitative analysis using inductive thematic analysis with consensual validation approach. Open coding followed by iterative refinement through team discussion. Five researchers involved in iterative coding and theme development. No formal inter-rater reliability coefficient computed for qualitative coding.

Main result

The study found that human-LLM disagreement results from recurring, identifiable causes: "The most frequent sources of divergence involved missing explicit terminology in abstracts, conceptual boundary ambiguity, and models over-relying on surface lexical cues." Analysis of six software engineering systematic reviews with over 1,000 papers revealed Cohen's Kappa values "ranging from 0.52 to 0.77," indicating moderate to substantial agreement. The researchers identified seven disagreement patterns: Term-Boundary, Abstract-Information Omission, LLM-Keyword Overweight, LLM-Not Main Focus, LLM-Incorrect Topic Inference, Human-Error, and Operationalization-Criteria Combination.

Reports effect sizes.

Research paradigm

Interpretivist/Qualitative

Author conclusions

The authors conclude that "LLMs have the potential to support researchers in title-abstract screening, which is a critical step in SRs. In this paper, we investigate how LLMs fail during screening. Through qualitative analysis of disagreements between LLMs and human raters across six software engineering SRs, we derived a taxonomy of seven disagreement patterns. Based on these patterns, we proposed five recommendations to help design more reliable LLM-assisted screening workflows. Our findings suggest that LLM failures result from recurring, identifiable causes, such as boundary ambiguity in key terms, information omission in abstracts, keyword overemphasization, and incorrect topic inference. Recognizing these patterns allows researchers to take targeted preventive measures, such as validating semantic understanding before deployment, running multiple LLMs, evaluating each criterion separately, and focusing validation efforts on borderline cases."

Risk of bias

Single rater bias in two ongoing studies (AI4T, AI4TInd); Reference standard bias: human consensus treated as ground truth despite human screening being subject to error; Selection bias: limited to software engineering domain only; Model version specificity: results reflect specific LLM versions available at time of study; Publication status bias: three unpublished/ongoing studies (GenAIEdu, AI4T, AI4TInd) had not undergone full peer review; Prompt design bias: zero-shot prompting only; different prompting strategies not tested; Reference standard bias - human consensus may contain errors and interpretation variability; Model-specific bias - results limited to specific LLM versions (gemini-2.5-flash, openai/gpt-4.1-mini, anthropic/claude-haiku-4.5); Selection bias - analysis limited to software engineering SRs, may not generalize to other domains; Prompting bias - zero-shot mode only, other prompting strategies not tested; Temporal bias - results reflect specific model versions available at time of study; Limited to zero-shot prompting mode—other prompting strategies not explored; Specific model versions used may not generalize as LLMs evolve; Three of six subject studies were unpublished/ongoing at time of analysis, not peer-reviewed; Human consensus used as reference standard, but human screening itself subject to error; Domain-specific to software engineering SRs; generalization to other domains unclear

Limitations

  • "This study has limitations that should be considered when interpreting the findings
  • First, our analysis is limited to the title-abstract screening stage of SRs
  • Disagreement patterns at other stages, such as full-text screening or data extraction, may differ and warrant separate investigation
  • Second, while four studies employed multiple independent researchers, the two ongoing studies relied on a single rater, which may introduce bias in the coding of disagreements..
  • Third, all LLM screening was conducted in zero-shot mode using specific models
  • Different prompting strategies, such as few-shot or chain-of-thought prompting, may yield different disagreement patterns and are not reflected in our taxonomy..

Open questions raised

  • Disagreement patterns at other screening stages (full-text screening, data extraction) warrant separate investigation
  • Impact of different prompting strategies (few-shot, chain-of-thought) on disagreement patterns not examined
  • Generalization beyond software engineering domain needed
  • Validation of recommended interventions (R1-R5) requires quantitative studies
  • Community efforts needed to develop normative guidelines on LLM usage in systematic reviews
  • Need for quantitative studies validating the impact of the proposed five recommendations
Data: https://dx.doi.org/10.5281/zenodo.20695824; Study data and analysis materials; Data Availability statementExtracted from: pdfAgreement 60%

Explore related topics

Related papers