12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How LLMs Cite and Why It Matters: A Cross-Model Audit of Reference Fabrication in AI-Assisted Academic Writing and Methods to Detect Phantom Citations

MZ Naser · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative experimental audit across ten commercially deployed LLMs.

Sample

N = 69557, 13 groups

Primary method

Chi-squared tests of independence with Cramér's V as effect size measure; bootstrap confidence intervals (10,000 resamples) at 95% level; Spearman's rank correlation (ρ) for sensitivity analysis; 5-fold stratified cross-validation for classifiers; leave-one-model-out (LOMO) cross-validation for generalization testing; fuzzy title matching using token-level Jaccard distance after text normalization; Pairwise Jaccard similarity for cross-model convergence; gradient-boosted machine (GBM), random forest (RF), and logistic regression classifiers with AUC, average precision (AP), and accuracy metrics; automated validation using GPT-4.1-mini with web search on stratified sample. Statistical software not explicitly named.

Main result

The study found that "citation hallucination seems to be an induced, not intrinsic, phenomenon" with "no model spontaneously generates formal citations when unprompted (0 of 3,030 responses)," and "hallucination rates range from 11.4% (GPT-5-mini) to 56.8% (haiku-4.5), a fivefold variation." Additionally, "prompts requesting 'recent and influential' references produced a hallucination rate of 74.1%, compared to 55.0% for prompts requesting 'seminal and foundational' references," and "among unique title strings cited by only one model, the match rate against our verification pipeline is 16.5%. For titles cited by two models, this rises to 87.4%. At three or more models, the match rate reaches 95.6%, a 5.8-fold improvement over the single-model baseline."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/Empiricist - quantitative measurement of hallucination rates across models using automated verification pipeline and statistical inference

Author conclusions

"This study presents one of the largest multi-model citation-hallucination audits conducted to date, spanning 10 commercially deployed LLMs, 4 academic domains, 2 temporal framings, 3 independent replications, and an unprompted control condition." The authors conclude: "Citation hallucination seems to be an induced, not intrinsic, phenomenon. This is due to the observation wherein no model spontaneously generates formal citations when unprompted (0 of 3,030 responses), and all observed fabrication is attributable to the explicit request to cite. This finding reframes hallucination mitigation as a problem of managing specific prompt-response interactions rather than addressing a deep generative tendency." Additionally, "citation hallucination can be filterable. More specifically, multi-model consensus (where three or more LLMs agree on a citation) raises accuracy from 16.5 to 95.6%. Similarly, within-model repetition (a citation recurring across replications) raises accuracy from 28.6 to 88.9%." Finally, "improvement is provider-specific. For example, OpenAI reduced hallucination by 33.9% between the examined model generations. In contrast, Anthropic's hallucination rate increased by 8.0% over the same interval."

Risk of bias

Model selection bias: Ten models selected for comparability, excluding reasoning-specialized and frontier models; Domain selection bias: Four domains selected to span contrasting literature volumes; may not represent heterogeneity across law, humanities, social sciences, or non-English-language scholarship; Verification pipeline bias: Conservative estimates with ~10.7% false-negative rate in hallucination classification; Training data composition differences across models not controlled; Temporal bias in training data cutoffs across models unknown; Verification pipeline false-negative rate: 10.7% of citations classified as hallucinated were found to be real upon independent validation; Model selection bias: frontier reasoning models (o3, DeepSeek-R1) and very large models excluded; Domain selection bias: limited to 4 academic domains, excluding law, humanities, many social science subfields, and non-English-language scholarship; Verification threshold arbitrariness: Rank ordering partially but not completely preserved under inclusive threshold (65-79 score range); Selection bias in model selection: intentionally selected mid-tier and smaller models, excluding frontier reasoning-specialized systems and very large models; Selection bias in domain coverage: limited to four academic domains (structural engineering, climate/environmental science, biomedical research, NLP/AI), excluding law, humanities, social sciences, and non-English-language scholarship; Measurement bias: verification pipeline false-negative rate of 10.7% in hallucinated stratum suggests conservative hallucination rate estimates; Temporal bias: model performance may vary over time as training data and alignment methods evolve; Language bias: study conducted entirely in English, excluding non-English academic publication patterns; Access bias: systematic bias toward open-access works in model outputs (77.3%-91.8% OA citation rate vs. 50-55% global baseline); Publisher bias: models show concentration in papers from specific publishers and high-citation venues

Limitations

  • "This study evaluates unaugmented text-generation models and therefore reflects citation behavior when the model must generate references from parametric memory alone
  • Systems that incorporate retrieval-augmented generation (RAG) and query curated bibliographic indices can ground outputs in external records and are expected to exhibit different error modes and typically lower phantom-reference rates." Additionally, "the scope of inference is also shaped by the model set, which was intentionally selected for comparability across providers (vs
  • exhaustive coverage of all available offerings)
  • Reasoning-specialized models (e.g., o3 and DeepSeek-R1) and very large frontier systems (e.g., GPT-5.2 and Claude Opus) were excluded because these models may rely on different internal deliberation policies, compute allocation, cost, or tuning objectives." Furthermore, "Interpretation is also conditioned on the limited set of academic domains used to elicit citation behavior, since disciplinary ecosystems differ in data size, venue structure, indexing coverage, and citation conventions." Lastly, "the verification pipeline, while cross-validated against an independent automated check, remains an operational classifier that can mislabel a minority of records under ambiguity or sparse metadata
  • The observed recovery of true citations (about 10%) within the hallucinated stratum implies that some legitimate references are being classified as phantom under the current thresholds, which makes the reported hallucination rates conservative upper bounds under this measurement definition."

Open questions raised

  • Need for extension to additional academic domains (law, humanities, social sciences, non-English-language scholarship)
  • Evaluation of reasoning-specialized and frontier models excluded from current study
  • Comparison of citation behavior in retrieval-augmented generation (RAG) systems
  • Expert adjudication on ambiguous citation verification cases to strengthen calibration
  • Investigation of specific alignment and training data choices driving developer-specific differences
  • Understanding mechanisms explaining Anthropic Haiku regression despite newer generation
Data: The authors state: "The classifier, trained model, and the full citation dataset are publicly released to support replication and downstream tool development." Specific URLs are not provided in the paper text.; Full citation dataset: 69,557 citation instances with verification labels (released publicly per data availability statement); Stratified validation sample: 225 citations for independent verification; "The classifier, trained model, and the full citation dataset are publicly released to support replication and downstream tool development." Specific data availability link not provided in document, but authors state release is available (see Data availability statement referenced in text). Dataset comprises 69,557 citation instances extracted from 15,150 model responses.Code: The paper mentions: "The trained model and feature extraction code are released alongside this paper (see Data availability statement)" but no specific GitHub or GitLab URLs are provided in the text.; Trained gradient-boosted machine learning classifier and feature extraction code released (location per Data availability statement in paper); "The trained model and feature extraction code are released alongside this paper (see Data availability statement)." No specific GitHub or repository URL provided in text.Extracted from: pdfAgreement 48%

Explore related topics

Related papers