How LLMs Cite and Why It Matters: A Cross-Model Audit of Reference Fabrication in AI-Assisted Academic Writing and Methods to Detect Phantom Citations
MZ Naser · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative experimental audit across ten commercially deployed LLMs.
Sample
N = 69557, 13 groups
Primary method
Chi-squared tests of independence with Cramér's V as effect size measure; bootstrap confidence intervals (10,000 resamples) at 95% level; Spearman's rank correlation (ρ) for sensitivity analysis; 5-fold stratified cross-validation for classifiers; leave-one-model-out (LOMO) cross-validation for generalization testing; fuzzy title matching using token-level Jaccard distance after text normalization; Pairwise Jaccard similarity for cross-model convergence; gradient-boosted machine (GBM), random forest (RF), and logistic regression classifiers with AUC, average precision (AP), and accuracy metrics; automated validation using GPT-4.1-mini with web search on stratified sample. Statistical software not explicitly named.
Main result
The study found that "citation hallucination seems to be an induced, not intrinsic, phenomenon" with "no model spontaneously generates formal citations when unprompted (0 of 3,030 responses)," and "hallucination rates range from 11.4% (GPT-5-mini) to 56.8% (haiku-4.5), a fivefold variation." Additionally, "prompts requesting 'recent and influential' references produced a hallucination rate of 74.1%, compared to 55.0% for prompts requesting 'seminal and foundational' references," and "among unique title strings cited by only one model, the match rate against our verification pipeline is 16.5%. For titles cited by two models, this rises to 87.4%. At three or more models, the match rate reaches 95.6%, a 5.8-fold improvement over the single-model baseline."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Empiricist - quantitative measurement of hallucination rates across models using automated verification pipeline and statistical inference
Author conclusions
"This study presents one of the largest multi-model citation-hallucination audits conducted to date, spanning 10 commercially deployed LLMs, 4 academic domains, 2 temporal framings, 3 independent replications, and an unprompted control condition." The authors conclude: "Citation hallucination seems to be an induced, not intrinsic, phenomenon. This is due to the observation wherein no model spontaneously generates formal citations when unprompted (0 of 3,030 responses), and all observed fabrication is attributable to the explicit request to cite. This finding reframes hallucination mitigation as a problem of managing specific prompt-response interactions rather than addressing a deep generative tendency." Additionally, "citation hallucination can be filterable. More specifically, multi-model consensus (where three or more LLMs agree on a citation) raises accuracy from 16.5 to 95.6%. Similarly, within-model repetition (a citation recurring across replications) raises accuracy from 28.6 to 88.9%." Finally, "improvement is provider-specific. For example, OpenAI reduced hallucination by 33.9% between the examined model generations. In contrast, Anthropic's hallucination rate increased by 8.0% over the same interval."
Risk of bias
Model selection bias: Ten models selected for comparability, excluding reasoning-specialized and frontier models; Domain selection bias: Four domains selected to span contrasting literature volumes; may not represent heterogeneity across law, humanities, social sciences, or non-English-language scholarship; Verification pipeline bias: Conservative estimates with ~10.7% false-negative rate in hallucination classification; Training data composition differences across models not controlled; Temporal bias in training data cutoffs across models unknown; Verification pipeline false-negative rate: 10.7% of citations classified as hallucinated were found to be real upon independent validation; Model selection bias: frontier reasoning models (o3, DeepSeek-R1) and very large models excluded; Domain selection bias: limited to 4 academic domains, excluding law, humanities, many social science subfields, and non-English-language scholarship; Verification threshold arbitrariness: Rank ordering partially but not completely preserved under inclusive threshold (65-79 score range); Selection bias in model selection: intentionally selected mid-tier and smaller models, excluding frontier reasoning-specialized systems and very large models; Selection bias in domain coverage: limited to four academic domains (structural engineering, climate/environmental science, biomedical research, NLP/AI), excluding law, humanities, social sciences, and non-English-language scholarship; Measurement bias: verification pipeline false-negative rate of 10.7% in hallucinated stratum suggests conservative hallucination rate estimates; Temporal bias: model performance may vary over time as training data and alignment methods evolve; Language bias: study conducted entirely in English, excluding non-English academic publication patterns; Access bias: systematic bias toward open-access works in model outputs (77.3%-91.8% OA citation rate vs. 50-55% global baseline); Publisher bias: models show concentration in papers from specific publishers and high-citation venues
Limitations
- "This study evaluates unaugmented text-generation models and therefore reflects citation behavior when the model must generate references from parametric memory alone
- Systems that incorporate retrieval-augmented generation (RAG) and query curated bibliographic indices can ground outputs in external records and are expected to exhibit different error modes and typically lower phantom-reference rates." Additionally, "the scope of inference is also shaped by the model set, which was intentionally selected for comparability across providers (vs
- exhaustive coverage of all available offerings)
- Reasoning-specialized models (e.g., o3 and DeepSeek-R1) and very large frontier systems (e.g., GPT-5.2 and Claude Opus) were excluded because these models may rely on different internal deliberation policies, compute allocation, cost, or tuning objectives." Furthermore, "Interpretation is also conditioned on the limited set of academic domains used to elicit citation behavior, since disciplinary ecosystems differ in data size, venue structure, indexing coverage, and citation conventions." Lastly, "the verification pipeline, while cross-validated against an independent automated check, remains an operational classifier that can mislabel a minority of records under ambiguity or sparse metadata
- The observed recovery of true citations (about 10%) within the hallucinated stratum implies that some legitimate references are being classified as phantom under the current thresholds, which makes the reported hallucination rates conservative upper bounds under this measurement definition."
Open questions raised
- Need for extension to additional academic domains (law, humanities, social sciences, non-English-language scholarship)
- Evaluation of reasoning-specialized and frontier models excluded from current study
- Comparison of citation behavior in retrieval-augmented generation (RAG) systems
- Expert adjudication on ambiguous citation verification cases to strengthen calibration
- Investigation of specific alignment and training data choices driving developer-specific differences
- Understanding mechanisms explaining Anthropic Haiku regression despite newer generation
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations