Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
Mingzhe Li, Zhiqiang Lin, Shiqing Ma · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-agent framework with four cascading modules: Reference Extractor (vision-LLM based PDF parsing), Cascading Evidence Collector (eight bibliographic connectors), Field Matcher (deterministic rule matching), and Class-specialist Judgers.
Primary method
Design science with iterative evaluation and ablation studies
Main result
CITETRACER achieves high performance on citation hallucination detection. The system "reaches 97.1% accuracy on the synthetic benchmark, with class-level F1 of 97.0, 95.8, and 98.5 for REAL, POTENTIAL, and HALLUCINATED, respectively, and detects 97.1% of fabrications on the real-world set without abstaining."
Research paradigm
empirical evaluation of computational artifact
Author conclusions
"We reframed citation hallucination detection from a binary found-or-not problem into a 12-code taxonomy and built a four-module cascading multi-agent detector that follows the taxonomy's structure: a deterministic rule matcher closes VALID and HALLUCINATED cases at near-zero cost, an ordered cascade over eight bibliographic connectors collects evidence before any LLM call, and three specialist agents adjudicate disjoint taxonomy slices with calibrated evidence thresholds." CITETRACER reaches 97.1% accuracy on synthetic data and 97.1% recall on real-world data.
Risk of bias
Domain bias: evaluation concentrated on CS/ML literature; Coverage bias: limited to citations indexed in queried bibliographic sources; API rate-limiting could produce false negatives in high-concurrency settings; Domain bias: Evaluation limited to computer science papers, especially ML literature; Selection bias: Synthetic benchmark built from only 50 recent ML and CS papers; Source coverage bias: Dependence on availability in specific bibliographic connectors (DBLP, Crossref, ACL, arXiv, OpenAlex, Semantic Scholar, Europe PMC, PubMed); API rate limiting: External dependencies may fail under high concurrency
Limitations
- The authors state that "our evaluation concentrates on Computer Science papers, especially the ML literature
- on citations from other fields with less standard formats, more complex structures, or limited coverage in the bibliographic connectors we query, the pipeline may miss candidates and emit incorrect HALLUCINATED verdicts." They also note that "under high-concurrency verification, parallel calls to the eight Scholar Connectors can trigger API rate limits and drop candidate evidence."
Open questions raised
- Fine-grained taxonomy for citation hallucination detection beyond binary verdicts
- Field-level audit capability for citation verification
- Robust PDF parsing for reference extraction
- Comprehensive retrieval pipeline covering diverse bibliographic sources
- Extension to non-CS domains with different citation formats and coverage
- The authors identify the need for improved handling of citations from fields beyond Computer Science with non-standard formats. They also mention the need for a future Scholar Connector router that routes each citation to the most appropriate connector by venue, publisher, and documented API coverage to reduce API rate-limiting issues.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations