12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences

Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic document analysis combining OCR-based citation extraction, database matching using fuzzy string similarity (Levenshtein distance with threshold 0.9), and manual verification of candidate HalluCitations.

Sample

N = 17000, 12 groups

Primary method

Descriptive statistics (mean, standard deviation, quartiles for citation counts); character-level fuzzy matching using normalized Levenshtein distance with similarity threshold of 0.9; proportional analysis (percentages of HalluCited papers by venue and track); frequency analysis (hit rates for papers containing multiple HalluCitation candidates). The study reports cumulative statistics and frequency distributions across bins of candidate counts.

Main result

The study found that "Almost 300 papers contain at least one HalluCitation, mostly published in 2025. Notably, half of these papers were identified at EMNLP 2025, the most recent conference, indicating that this issue is rapidly increasing. Moreover, over 100 such papers are accepted as main and Findings papers in EMNLP 2025, affecting the credibility." Additionally, "when a paper contains around four HalluCitation candidates, nearly three out of four papers indeed include HalluCitation," demonstrating a practical detection threshold.

Reports effect sizes.

Research paradigm

empirical-positivist

Author conclusions

The authors conclude: "We investigated the presence of HalluCitations in accepted papers at ACL conferences by examining all papers published at ACL, NAACL, and EMNLP in 2024 and 2025. As a result, we identified nearly 300 HalluCited papers, more than half of which appeared at EMNLP 2025." They emphasize that "the presence of HalluCitations should not be treated as grounds for immediate penalties" and argue "it is important to introduce dedicated HalluCitation detection tools" and "integrate them into existing toolkits such as ACL pubcheck." They recommend: "It is also worth noting that the HalluCited papers were accepted due to their substantive contributions. Accordingly, authors of HalluCited papers should not be penalized post hoc."

Risk of bias

Selection bias: analysis limited to accepted papers; rejected papers not analyzed; Detection bias: OCR and parsing errors may introduce false positives or negatives; Threshold bias: fuzzy matching similarity threshold of 0.9 may miss some citations or include spurious matches; Incomplete disclosure: preprint opt-in rate around 20%, unclear if HalluCited papers preferentially opt in or out; Scope bias: analysis restricted to six recent conferences; earlier years and other venues not analyzed; OCR parsing errors may introduce noise in citation extraction; Database matching threshold (0.9) may miss valid citations with title variations; Manual verification bias: authors may apply conservative verification criteria inconsistently; Limited to papers available in ACL Anthology and arXiv; other venues not analyzed; Preprint opt-in disclosure bias in ARR analysis (~20% disclosure rate may not be representative); Selection bias: analysis limited to accepted papers rather than full submission cohort; Selection bias: Only accepted papers analyzed; rejected papers not available. Detection bias: OCR errors and fuzzy matching thresholds (0.9 similarity) may introduce noise. Researcher bias: Manual verification conducted only by paper authors without inter-rater reliability assessment. Temporal bias: Analysis limited to 2024-2025; earlier years not representative.

Limitations

  • "This study focuses on six recent top-tier NLP conferences
  • Although the scope could be expanded, we limit our analysis for several reasons." The authors note "our methodology prioritizes precision, and the reported results should be interpreted as a lower bound
  • Additional HalluCitations may exist beyond those detected here." They also state "The transparency of the peer-review process remains limited, and a more thorough analysis is currently infeasible due to restricted access to review materials
  • Moreover, rejected papers are generally not publicly available." Additionally, "we employed MinerU for OCR, although alternative OCR tools and additional engineering improvements are possible," and they "focused on arXiv and the ACL Anthology as solid and well-established data sources."

Open questions raised

  • Expansion to other AI/ML conferences beyond top-tier NLP venues
  • Analysis of rejected papers (currently not publicly available)
  • Broader investigation of HalluCitations in journals and other scholarly publication venues
  • Improved OCR and parsing methods for more comprehensive coverage
  • Greater transparency in peer-review processes and archiving of review materials
  • Development and validation of automated HalluCitation detection systems for integration into author toolkits
Data: ACL Anthology papers (17,000+ papers from 2024-2025); arXiv papers referenced in ACL Anthology; ACL Rolling Review (ARR) metadata via ARR Statistics Dashboard; DBLP bibliographic database; OpenAlex bibliographic database; ACL Anthology (publicly available at https://aclanthology.org/); arXiv (publicly available, accessed via Kaggle at https://www.kaggle.com/datasets/Cornell-University/arxiv/); DBLP XML dump (https://drops.dagstuhl.de/entities/artifact/10.4230/dblp.xml.2025-12-01); OpenAlex (publicly available via REST APIs); ACL Rolling Review (ARR) public dashboard (https://stats.aclrollingreview.org/); List of HalluCited papers provided in Appendix B; ACL Anthology (https://aclanthology.org/); arXiv (https://arxiv.org/); OpenAlex (via REST API PyAlex); Google Scholar (https://scholar.google.com/); Semantic Scholar (https://www.semanticscholar.org/); ACL Rolling Review (ARR) Statistics Dashboard (https://stats.aclrollingreview.org/); OpenReview dataCode: MinerU (Wang et al., 2024a) - OCR-based reference extraction; GROBID (GRO, 2008-2025) - Bibliographic reference parsing; RapidFuzz library (Bachmann, 2025) - Fuzzy matching; Ribiber - Citation normalization tool referenced (https://github.com/yuchenlin/rebiber)Extracted from: pdfAgreement 56%

Explore related topics

Related papers