HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic document analysis combining OCR-based citation extraction, database matching using fuzzy string similarity (Levenshtein distance with threshold 0.9), and manual verification of candidate HalluCitations.
Sample
N = 17000, 12 groups
Primary method
Descriptive statistics (mean, standard deviation, quartiles for citation counts); character-level fuzzy matching using normalized Levenshtein distance with similarity threshold of 0.9; proportional analysis (percentages of HalluCited papers by venue and track); frequency analysis (hit rates for papers containing multiple HalluCitation candidates). The study reports cumulative statistics and frequency distributions across bins of candidate counts.
Main result
The study found that "Almost 300 papers contain at least one HalluCitation, mostly published in 2025. Notably, half of these papers were identified at EMNLP 2025, the most recent conference, indicating that this issue is rapidly increasing. Moreover, over 100 such papers are accepted as main and Findings papers in EMNLP 2025, affecting the credibility." Additionally, "when a paper contains around four HalluCitation candidates, nearly three out of four papers indeed include HalluCitation," demonstrating a practical detection threshold.
Reports effect sizes.
Research paradigm
empirical-positivist
Author conclusions
The authors conclude: "We investigated the presence of HalluCitations in accepted papers at ACL conferences by examining all papers published at ACL, NAACL, and EMNLP in 2024 and 2025. As a result, we identified nearly 300 HalluCited papers, more than half of which appeared at EMNLP 2025." They emphasize that "the presence of HalluCitations should not be treated as grounds for immediate penalties" and argue "it is important to introduce dedicated HalluCitation detection tools" and "integrate them into existing toolkits such as ACL pubcheck." They recommend: "It is also worth noting that the HalluCited papers were accepted due to their substantive contributions. Accordingly, authors of HalluCited papers should not be penalized post hoc."
Risk of bias
Selection bias: analysis limited to accepted papers; rejected papers not analyzed; Detection bias: OCR and parsing errors may introduce false positives or negatives; Threshold bias: fuzzy matching similarity threshold of 0.9 may miss some citations or include spurious matches; Incomplete disclosure: preprint opt-in rate around 20%, unclear if HalluCited papers preferentially opt in or out; Scope bias: analysis restricted to six recent conferences; earlier years and other venues not analyzed; OCR parsing errors may introduce noise in citation extraction; Database matching threshold (0.9) may miss valid citations with title variations; Manual verification bias: authors may apply conservative verification criteria inconsistently; Limited to papers available in ACL Anthology and arXiv; other venues not analyzed; Preprint opt-in disclosure bias in ARR analysis (~20% disclosure rate may not be representative); Selection bias: analysis limited to accepted papers rather than full submission cohort; Selection bias: Only accepted papers analyzed; rejected papers not available. Detection bias: OCR errors and fuzzy matching thresholds (0.9 similarity) may introduce noise. Researcher bias: Manual verification conducted only by paper authors without inter-rater reliability assessment. Temporal bias: Analysis limited to 2024-2025; earlier years not representative.
Limitations
- "This study focuses on six recent top-tier NLP conferences
- Although the scope could be expanded, we limit our analysis for several reasons." The authors note "our methodology prioritizes precision, and the reported results should be interpreted as a lower bound
- Additional HalluCitations may exist beyond those detected here." They also state "The transparency of the peer-review process remains limited, and a more thorough analysis is currently infeasible due to restricted access to review materials
- Moreover, rejected papers are generally not publicly available." Additionally, "we employed MinerU for OCR, although alternative OCR tools and additional engineering improvements are possible," and they "focused on arXiv and the ACL Anthology as solid and well-established data sources."
Open questions raised
- Expansion to other AI/ML conferences beyond top-tier NLP venues
- Analysis of rejected papers (currently not publicly available)
- Broader investigation of HalluCitations in journals and other scholarly publication venues
- Improved OCR and parsing methods for more comprehensive coverage
- Greater transparency in peer-review processes and archiving of review materials
- Development and validation of automated HalluCitation detection systems for integration into author toolkits
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations