12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Fluency Without Fidelity: Errors in Citation-Attributed Claims in Large Language Model-Generated Literature Reviews in Mental Health

Jake Linardon, Mariel Messer, Cleo Anderson, Olivia Marie Soliman, Claudia Liu, Joseph Firth et al. · Journal of Technology in Behavioral Science · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s41347-026-00630-6

Methodology & findings

Study design

Systematic audit and content analysis of LLM-generated literature reviews.

Sample

N = 333, 7 groups

Primary method

Descriptive analyses only. Proportions of citation-claims assigned to each accuracy category (1-4) were calculated by dividing the number of citation-claims in each category by the total number of citation-claims evaluated for each condition. No inferential statistical tests were performed. Interrater reliability was assessed using Cohen's kappa (κ > 0.74). Data visualization used bar charts stratified by disorder and treatment modality. Independent evaluation by two raters with consensus resolution and third-party consultation when needed.

Main result

Across all conditions combined, citation-claims generated by the LLM were most often scored as "completely accurate" (N = 145; 43.5%). Minor inaccuracies were found in 18.3% of citation-claims. The remaining 38.2% of citation-claims scored 2 or below, found to either contained major inaccuracies (18.3%) or be contradictory/fabricated (19.8%). The study found that "only 43.5% were completely accurate, while 18.3%" had minor inaccuracies, with "fabrication rates exceeding 20% in many instances, and substantially exceeding those observed in human-authored research."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-positivist

Author conclusions

"This study provides clear evidence that LLM-generated literature reviews in mental health research are frequently undermined by serious citation-claim inaccuracies. The model typically failed to accurately align claims with the content of the sources it cited, with fabrication rates exceeding 20% in many instances, and substantially exceeding those observed in human-authored research. Citation-claim errors generated by LLMs may be becoming less apparent and harder to detect, as newer models increasingly cite real, verifiable sources while continuing to inaccurately interpret or represent the underlying evidence. Until fidelity is markedly improved and safeguards are in place, LLMs do not appear to be in a position to generate academic literature reviews in health-related domains where inaccuracies may carry significant adverse consequences."

Risk of bias

Single LLM model tested (ChatGPT-5 only) - generalizability to other models unclear; Single prompting strategy - accuracy may vary with different prompt variations; Limited to mental health domain - findings may not generalize to other disciplines; Subjective scoring framework (custom-developed scale) - potential for scoring bias despite high interrater reliability (κ > 0.74); Small sample of citation-claims per condition (≤21 per condition) - limits statistical inference and condition-specific conclusions; No blinding of raters to condition (disorder/modality) during evaluation; Conservative scoring approach (lowest score retained for multi-claim passages) may inflate error rates; Single LLM model (ChatGPT-5) tested; may not generalize to other models or versions; Single prompting strategy; alternative prompt formulations not tested; Discipline-specific limitation: mental health only; no cross-domain validation; Subjective accuracy scoring scale developed for this study despite interrater reliability checks (κ=0.74); Potential reviewer expectation bias despite consensus-based resolution and third-party arbitration; Single evaluator bias partially mitigated through dual-rater system with κ = 0.74 interrater reliability; Selection bias in choice of single LLM (ChatGPT-5) and specific prompting strategy; Potential observer bias in accuracy scoring despite use of structured criteria; Publication date bias in cited sources (42% published within last 6 years)

Limitations

  • The authors state: "First, the accuracy scoring scale was developed specifically for this study
  • Although the scoring criteria were informed by prior research examining quotation accuracy in human-authored research (Lazonder & Janssen, 2022), the absence of a universally accepted framework for evaluating LLM citation accuracy introduces a degree of subjectivity." Additionally: "Second, findings are limited to mental health research and may not generalize to other disciplines
  • Thus, inferences should be confined to mental health research, and future studies are needed to examine whether citation-claim accuracy varies systematically across fields." Third: "because this evaluation focused on a single LLM and a specific prompting strategy, the observed accuracy rates may not generalize to other models, versions, or prompt variations."

Open questions raised

  • Need for evaluation of citation-claim accuracy in other disciplines (beyond mental health)
  • Systematic comparison of citation-claim accuracy across different LLM models, versions, and prompting strategies
  • Investigation of how model architecture, training data, and prompt constraints affect accuracy rates
  • Evaluation of whether citation-claim accuracy varies systematically across academic fields
  • Development of universally accepted frameworks for evaluating LLM citation accuracy
  • Research on technical safeguards such as real-time source verification and citation grounding mechanisms
Data: "All relevant data are presented in the Supplementary Materials." Supplementary Tables 1-4 contain the 333 citation-claims evaluated, corresponding source publications, accuracy scores, and justifications. These are referenced as being available online at https://doi.org/10.1007/s41347-026-00630-6.; All relevant data are presented in the Supplementary Materials. Supplementary Tables 1-4 present the 333 citation-claims evaluated for each condition, the corresponding source publication, the assigned accuracy score for each citation-claim, and the justification for that score. Verbatim prompts are provided in the Supplementary Materials to facilitate replication.; "All relevant data are presented in the Supplementary Materials."Extracted from: pdfAgreement 51%

Explore related topics

Related papers