12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The Case of the Mysterious Citations

Amanda Bienz, Carl M. Pearson, Simon Garcia de Gonzalo · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Observational content analysis of academic conference proceedings.

Sample

> 1000, 14 groups

Primary method

Descriptive statistics only. The study reports: percentages of papers with citation errors (ranging from under 2% to nearly 6%), counts of citations with specific error types, and comparative analysis between 2021 and 2025 baseline and recent proceedings. No inferential statistics, hypothesis tests, or confidence intervals are reported.

Main result

The study found that "every 2025 proceeding we examined contains at least one mysterious citation" and "Both rephrased titles and mysterious citations were found in all four conference proceedings, with the fraction of papers in which they were found ranging from under 2% of papers to nearly 6%." Additionally, the authors report that "every mysterious citation identified in this study appears in published proceedings" and "these papers successfully passed peer review and editorial checks, yet still contain fabricated or corrupted references."

Reports effect sizes.

Research paradigm

Empirical positivism

Author conclusions

The authors conclude that "Our study provides the first systemic evidence that unverified, AI-hallucinated citations have begun to appear in peer-reviewed conferences proceedings. By comparing four major venues from 2021 and 2025, we show that every 2025 proceeding we examined contains at least one mysterious citation." They further state that "Combining individual diligence with explicit community-wide policy on this subject and developing automated tools could help improve traceability and maintain trust in the scientific process."

Risk of bias

Selection bias: Authors analyzed only four high-performance computing conferences with which they had familiarity as previous authors or program committee members; Confirmation bias: Unable to definitively prove citations were AI-generated without author acknowledgment; Ascertainment bias: Manual verification by authors may introduce subjective interpretation of 'mysterious citations'; Detection bias: Citation error detection may vary by availability and accessibility of external databases; Selection bias: Conferences selected based on authors' familiarity as previous authors or program committee members, not random selection; Verification bias: Manual verification by authors without blinding or inter-rater reliability checks; Attribution bias: Inability to definitively prove LLM generated citations despite correlation with AI adoption; Sample size variation: Conferences vary in size (less than 50 to over 100 publications), affecting comparability; Non-acknowledgment confound: Papers acknowledged only selective AI use (e.g., grammar editing) while mystery citations may have unacknowledged origins; Selection bias: Authors selected conferences based on their own familiarity rather than random sampling; Temporal confounding: Cannot definitively attribute citation errors to LLM use without author acknowledgment; Limited scope: Analysis restricted to HPC conferences, limiting generalizability; Verification bias: Manual validation by authors may introduce subjective judgment in error classification; Publication bias: Only examines published proceedings, not rejected papers or preprints

Limitations

  • The authors explicitly state: "It is not possible to prove that a citation error was generated through an LLM hallucination
  • While we are not aware of any confounding factors that would cause the incidence of human-derived bibliography errors to arise, we do not explicitly account for them either." Additionally, the study acknowledges that determining whether text was generated by LLM is "an unsolved problem, and may even be impossible," and the authors were unable to verify that AI was actually used to generate the incorrect citations found despite conference requirements for acknowledgment.

Open questions raised

  • The authors identify the need for: (1) development of automated citation verification tools; (2) explicit community policies defining unacceptable citation practices and penalties; (3) clearer policies on generative AI use disclosure in academic writing; (4) extension of findings beyond HPC conferences to other research domains; (5) investigation of whether improved generative AI models will reduce hallucinated citations.
  • The authors identify the need for: (1) development of citation-verification tools made publicly available to authors and reviewers; (2) explicit community-wide policies defining unacceptable citation practices and associated penalties; (3) automated tools to improve traceability; (4) investigation across research domains beyond high-performance computing; (5) evaluation of whether research papers can remain reliable methods for sharing scientific progress given increased AI use.
  • Development of detection tools: Authors note that citation-verification tools like GPTZero's hallucination check exist but are "not publicly available to authors or reviewers"
  • Understanding prevalence across disciplines: While the study focuses on HPC conferences, authors note that "AI-generated citations are appearing across research domains"
  • Distinguishing legitimate LLM use from misuse: Authors acknowledge that "deciding whether a given piece of text is generated by an LLM used by an adversarial author is an unsolved problem, and may even be impossible"
  • Community policy development: Need for explicit policies defining unacceptable citation practices and associated penalties
Extracted from: pdfAgreement 59%

Explore related topics

Related papers