12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs de Vaan, Paul Ginsparg, Yian Yin · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale audit of 111 million references across 2.5 million papers in arXiv, bioRxiv, SSRN, and PubMed Central.

Primary method

Design science; systematic audit and measurement framework development

Main result

The study found a "sharp rise in non-existent references following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone." Hallucinated citations are "diffusely embedded across many papers but especially pronounced in fields with rapid AI uptake, in manuscripts with linguistic signatures of AI-assisted writing, and among small and early-career author teams." The analysis shows that "hallucinated references disproportionately assign credit to already prominent and male scholars, suggesting that LLM-generated errors may reinforce existing inequities in scientific recognition."

Research paradigm

Positivist/Empiricist

Author conclusions

The authors conclude that "hallucinated citations are widespread, disproportionately produced by junior researchers, and they channel credit toward established scholars." They state that "Science is arguably the best-case domain for hallucination detection: it has rich indexing systems, large-scale bibliographic databases, strong norms of citation, layered editorial review, and now a rapidly developing toolkit of automated verifiers. Yet even here, hallucinated content enters the record, persists from preprint through publication, and compounds through path-dependent citation networks and model training loops." Finally, they emphasize that "Science is the most measurable instance of what is likely a far broader phenomenon: the infiltration of AI-hallucinated content into the knowledge systems on which modern institutions make consequential decisions."

Risk of bias

Selection bias: The study focuses on four specific corpora (arXiv, bioRxiv, SSRN, PMC) which may not represent all scientific literature; Detection bias: The methodology relies on citation matching algorithms which may have different accuracy rates across fields and citation formats; Measurement bias: The baseline error rate (pre-LLM unmatched citations) may not be uniform across all datasets and time periods; Coverage bias: Limited coverage for fields whose citation norms omit titles; False positive/negative risk: LLM-based cleaning step (GPT-4o-mini) could reintroduce model biases; Selection bias: Study restricted to four major preprint/publication platforms with systematic differences in fields and citation-parsing techniques; Detection bias: Pipeline may miss hallucinations in non-English sources, niche venues, and heavily mathematical documents; Confounding: Rising unmatched rates could reflect changes in citation practices, database coverage, or indexing rather than LLM hallucination; Temporal confounding: Other factors besides LLM adoption could explain post-2023 citation patterns; Detection bias: False negatives from non-academic sources and imperfect parsing; false positives from real titles with incorrect metadata; Database coverage bias: Limited indexing in Semantic Scholar and OpenAlex for non-mainstream venues; Field-specific bias: Limited coverage for mathematical fields and niche venues with different citation norms; Selection bias: Focus on preprint servers and specific peer-reviewed journals may not represent all scientific literature; Temporal bias: Analysis extends only through August 2025; may not capture full trajectory

Limitations

  • The authors acknowledge that "the pipeline may produce false negatives for niche venues and heavily mathematical documents, and it has limited coverage for fields whose citation norms omit titles
  • False positives may arise when LLMs generate a real title with incorrect metadata, though existing estimates suggest such cases are relatively minor compared with the magnitudes we report." Additionally, "hallucinated titles represent the most detectable form of the problem
  • The more prevalent and harder-to-detect variant—real citations deployed to support claims the cited references do not actually make—remains an open challenge for which reliable detection methods remain under active development."

Open questions raised

  • The need to develop reliable detection methods for the harder-to-detect variant: real citations deployed to support claims the cited references do not actually make
  • Understanding hallucinations in domains that lack citation infrastructure (government reports, legal filings, clinical documentation, corporate knowledge bases, journalism)
  • Detecting hallucinated assertions embedded in unstructured prose beyond citation titles
  • Understanding long-term effects of hallucinated citations on scientific knowledge accumulation
  • Determining appropriate regulatory and policy responses to LLM hallucinations in knowledge work
  • Detection of real citations deployed to support claims the cited references do not actually make (more prevalent but harder-to-detect variant)
Data: arXiv (1,465,145 preprints, Jan 2020–Aug 2025); bioRxiv (261,928 preprints); SSRN (421,698 preprints); PubMed Central (10% random sample of 374,807 manuscripts between 2020 and 2025); PubMed Central (374,807 manuscripts from 10% random sample, 2020-2025); arXiv (1,465,145 preprints, Jan 2020–Aug 2025): reference extraction from LaTeX source files and GROBID processing of PDFs; bioRxiv (261,928 preprints): XML-formatted citation records; SSRN (421,698 preprints): metadata via Crossref API; PubMed Central (374,807 manuscripts, 10% random sample, 2020–2025): full-text reference extractionCode: Not mentioned; no GitHub or code repository links providedExtracted from: pdfAgreement 67%

Explore related topics

Related papers