LLM hallucinations in the wild: Large-scale evidence from non-existent citations
Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs de Vaan, Paul Ginsparg, Yian Yin · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale audit of 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central.
Primary method
Design science; systematic audit and measurement framework development
Main result
The study found a "sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone." Hallucinated citations are "diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams." The analysis shows that "hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition."
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude that "hallucinated citations are widespread, disproportionately produced by junior researchers, and they channel credit toward established scholars." They state that "Science is arguably the best-case domain for hallucination detection: it has rich indexing systems, large-scale bibliographic databases, strong norms of citation, layered editorial review, and now a rapidly developing toolkit of automated verifiers. Yet even here, hallucinated content enters the record, persists from preprint through publication, and compounds through path-dependent citation networks and model training loops." Finally, they emphasize that "Science is the most measurable instance of what is likely a far broader phenomenon: the infiltration of AI-hallucinated content into the knowledge systems on which modern institutions make consequential decisions."
Risk of bias
Selection bias: The study focuses on four specific corpora (arXiv, bioRxiv, SSRN, PMC) which may not represent all scientific literature; Detection bias: The methodology relies on citation matching algorithms which may have different accuracy rates across fields and citation formats; Measurement bias: The baseline error rate (pre-LLM unmatched citations) may not be uniform across all datasets and time periods; Coverage bias: Limited coverage for fields whose citation norms omit titles; False positive/negative risk: LLM-based cleaning step (GPT-4o-mini) could reintroduce model biases; Selection bias: Study restricted to four major preprint/publication platforms with systematic differences in fields and citation-parsing techniques; Detection bias: Pipeline may miss hallucinations in non-English sources, niche venues, and heavily mathematical documents; Confounding: Rising unmatched rates could reflect changes in citation practices, database coverage, or indexing rather than LLM hallucination; Temporal confounding: Other factors besides LLM adoption could explain post-2023 citation patterns; Detection bias: False negatives from non-academic sources and imperfect parsing; false positives from real titles with incorrect metadata; Database coverage bias: Limited indexing in Semantic Scholar and OpenAlex for non-mainstream venues; Field-specific bias: Limited coverage for mathematical fields and niche venues with different citation norms; Selection bias: Focus on preprint servers and specific peer-reviewed journals may not represent all scientific literature; Temporal bias: Analysis extends only through August 2025; may not capture full trajectory
Limitations
- The authors acknowledge that "the pipeline may produce false negatives for niche venues and heavily mathematical documents, and it has limited coverage for fields whose citation norms omit titles
- False positives may arise when LLMs generate a real title with incorrect metadata, though existing estimates suggest such cases are relatively minor compared with the magnitudes we report." Additionally, "hallucinated titles represent the most detectable form of the problem
- The more prevalent and harder-to-detect variant—real citations deployed to support claims the cited references do not actually make—remains an open challenge for which reliable detection methods remain under active development."
Open questions raised
- The need to develop reliable detection methods for the harder-to-detect variant: real citations deployed to support claims the cited references do not actually make
- Understanding hallucinations in domains that lack citation infrastructure (government reports, legal filings, clinical documentation, corporate knowledge bases, journalism)
- Detecting hallucinated assertions embedded in unstructured prose beyond citation titles
- Understanding long-term effects of hallucinated citations on scientific knowledge accumulation
- Determining appropriate regulatory and policy responses to LLM hallucinations in knowledge work
- Detection of real citations deployed to support claims the cited references do not actually make (more prevalent but harder-to-detect variant)
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations