How unique are hallucinated citations offered by generative Artificial Intelligence models?
Dirk HR Spennemann · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods approach combining: (1) systematic bibliographic search in Google Scholar and Google for the phantom citation 'Education Governance and Datafication' with extraction and analysis of 137 accessible source papers; (2) structured interrogation of ChatGPT 5-mini via three conversation protocols examining citation generation processes; (3) experimental generation of ten AI-written essays with citation analysis.
Sample
N = 137, 13 groups
Primary method
Descriptive statistics (frequency analysis, percentages, correlation analysis); digital forensic analysis (Wayback Machine verification); systematic database searches; pattern analysis of citation components; probability calculation (p=0.055 for duplicate reference generation without pagination; p=0.000675 with pagination).
Main result
The study found that "hallucinated citations are not random inventions but patterned recombinations of real authors, journals, dates, and keywords, with duplication occurring in nearly 30% of cases." Additionally, "while most references were genuine or partly accurate, 9.2% remained hallucinated, including an exact match to the most common phantom citation" in AI-generated essays.
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-positivist with qualitative analysis
Author conclusions
The authors conclude that "The hallucinated academic references created by ChatGPT are not random errors but predictable, pattern-driven artifacts of how genAI models generate text. These references are systematic reconstructions built from real authors, journals, and topical keywords." They further note that "ChatGPT-written essays will continue to carry hallucinated references even though the model has access to the web" and warn that "Given that citation laundering (i.e. citing a source because others cited it) is not uncommon, it may well be only a matter of time that such second-generation citations will occur."
Risk of bias
Selection bias: Study limited to papers found through Google Scholar and Google searches; Attribution bias: Authors cannot verify which AI models actually generated source papers; Temporal bias: Analysis conducted March 2026, potentially missing earlier or later instances; Journal selection bias: Focus on specific hallucinated title may not represent hallucination patterns across other fabricated citations; Selection bias: Papers identified through Google Scholar and Google searches may not represent the full population of publications containing the phantom citation; Attrition bias: Papers behind paywalls were excluded, potentially missing relevant data; Confirmation bias: The choice to focus on a single phantom citation may not generalize to other hallucinated references; Temporal bias: Analysis conducted at single timepoint (March 2026) does not capture longitudinal patterns; Source attribution bias: Assumption that ChatGPT is the generator without direct evidence from authors; Selection bias: Only 137 of 147 identified papers could be accessed; Platform bias: Reliance on Google Scholar and Google searches may miss papers indexed in other databases; Temporal bias: Data collection occurred March-31 2026, recent timeframe; Single model testing: Only ChatGPT 5-mini interrogated, not generalizable to other genAI models; Documentation bias: Authors did not attest to genAI use, only surmised ChatGPT involvement
Limitations
- The paper acknowledges that "a small number [of source papers] was behind non-institutional paywalls" limiting complete access, and notes that "at present, no empirical data exist to formally ascertain the existence and potential prevalence" of identical hallucinated references being regenerated across different runs
- Additionally, the study could only surmise rather than confirm ChatGPT as the source of hallucinations, as "none of the authors of the source papers attested to the use of genAi (let alone specific models) in the creation of their manuscripts."
Open questions raised
- No empirical data previously existed to formally ascertain the existence and prevalence of identical hallucinated references being regenerated in different AI runs
- Lack of clear explanation in literature regarding how genAI models arrive at hallucinated citations beyond basic concepts of genAI functionality
- Need for understanding why papers are back-dated to predate ChatGPT's public release
- Investigation of how hallucinated citations enter the formal academic citation universe through citation laundering
- Lack of empirical data on the prevalence of identical hallucinated references generated across different AI runs/chats
- Absence of clear explanation in literature regarding how genAI models arrive at hallucinated citations beyond basic functionality concepts
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations