12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025

Samar Ansari · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic content analysis and failure mode taxonomy development.

Sample

N = 100, 5 groups

Primary method

Descriptive frequency analysis. Percentage distributions reported for primary failure mode categories and secondary failure characteristics. Chi-square tests not reported. Statistical inference testing was not performed; the analysis is descriptive taxonomy-based classification.

Main result

The study found that "at least 53 of these papers (≈1%) contained fabricated citations that evaded detection throughout the review process" at NeurIPS 2025. A critical finding emerged: "every hallucination in our dataset exhibited compound failure modes, employing multiple deception strategies simultaneously." Specifically, "The distribution of secondary failure characteristics was dominated by Semantic Hallucination (63% of all citations) and Identifier Hijacking (29% of all citations)." The analysis reveals that "AI-generated hallucinations are predominantly Total Fabrications (66%), with citations invented wholesale rather than corrupted from real sources."

Reports effect sizes.

Research paradigm

Empirical-descriptive (systematic content analysis of documented artifacts)

Author conclusions

The authors conclude: "This study documents 100 AI-generated hallucinated citations that appeared in papers accepted by NeurIPS 2025, one of the world's most prestigious AI conferences. Each paper was reviewed by 3-5 expert researchers, yet fabricated citations evaded detection throughout the peer review process." They further state that "every single hallucination exhibited compound failure modes, with 63% layering semantic plausibility onto fabricated content and 29% incorporating identifier hijacking to create false verifiability." The authors emphasize that "the solution is implementable now. Automated citation verification tools exist and are already used in research integrity investigations," and recommend that "Conference organizers should mandate four-step verification: (1) existence check via web search and academic databases, (2) metadata consistency check confirming authors/title/venue match, (3) identifier validation confirming DOIs/arXiv IDs point to claimed papers, and (4) flagging semantically suspicious citations for human review."

Risk of bias

Selection bias: Only citations flagged by GPTZero's automated tool were analyzed; undetected hallucinations and citations with subtle misrepresentations would be missed; Detection tool bias: Classification depends entirely on GPTZero's detection methods and human verification process; Single conference bias: Data limited to NeurIPS 2025, potentially unrepresentative of other venues or disciplines; Coding bias: Manual classification by single author could introduce subjective judgment despite structured spreadsheet approach; Temporal bias: Analysis of single year snapshot may not capture evolving hallucination patterns; Selection bias: Only citations flagged by GPTZero's automated tool were analyzed; undetected hallucinations would be missed; Detection bias: Citations that appear correct but subtly misrepresent content would not be identified; Domain specificity bias: Data limited to single conference (NeurIPS 2025) and single discipline (AI/ML), limiting generalizability; Coder bias: Single author performed manual classification; no inter-rater reliability reported; Recall bias: Authors acknowledged uncertainty about whether analyzed fabrications are true de novo hallucinations or inherited from contaminated training data; Selection bias: Only citations flagged by GPTZero's automated tool were included; undetectable hallucinations are excluded; Classification bias: Author performs manual coding without inter-rater reliability assessment; Domain specificity bias: Data limited to AI/ML research at single conference (NeurIPS 2025); Detection tool limitations: GPTZero's tool may have systematic blind spots for certain fabrication types

Limitations

  • The authors acknowledge several limitations: "we analyzed only citations flagged by GPTZero's automated tool
  • Citations that appear correct but subtly misrepresent source content would not be detected." Second, "we cannot determine author intent, whether hallucinations resulted from deliberate fraud, negligent use of AI tools, or honest mistakes." Third, "our taxonomy classifies hallucinations by their most prominent failure characteristic as the primary code, with secondary characteristics captured separately
  • The boundaries between categories (particularly PAC vs
  • TF) can be ambiguous when citations blend multiple real and fake elements." Finally, "the data comes from a single conference (NeurIPS 2025) in a single discipline (AI/ML)." The authors note that "Different LLMs, different prompting strategies, or different research fields might produce different hallucination patterns."

Open questions raised

  • Distinguishing between pure fabrication and Contamination Inheritance (reproduction of pre-existing erroneous citations from contaminated training data)
  • Understanding the proportion of citations classified as Total Fabrication that are actually Contamination Inheritances
  • Developing training-data provenance tools to trace origins of hallucinated citations
  • Understanding whether fabrication strategies are evolving in response to detection tools across model generations
  • Longitudinal studies tracking hallucination patterns across model generations, training datasets, and research domains
  • Understanding the prevalence of Contamination Inheritance (where LLMs reproduce pre-existing errors from training data rather than fabricating de novo) versus pure fabrication across different models and domains
Data: GPTZero hallucinated citations dataset; GPTZero public report containing 100 verified hallucinations (full dataset referenced as [1] but specific URL not provided in text); GPTZero's systematic analysis of NeurIPS 2025 hallucinated citationsExtracted from: pdfAgreement 60%

Explore related topics

Related papers