CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text
Khashayar Khajavi, Shaghayegh Sadeghi, Rise Adhikari, Alexander Tessier · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Hybrid computational framework combining information retrieval, structured LLM-based verification, and threshold-based classification.
Main result
On the held-out test set, CITECHECK achieves 88.7 macro-F1 and 88.9% accuracy, outperforming GPT, Claude, and Gemini baselines, including web-search and few-shot variants. "CITECHECK improves over the strongest baseline by 5.8 macro-F1 points and 5.7 accuracy points." The results demonstrate that "reliable citation verification benefits from combining scholarly retrieval, structured LLM-based comparison, and calibrated decision rules."
Research paradigm
computational validation and empirical evaluation
Author conclusions
"We introduced CITECHECK, a framework for detecting hallucinated citations in scientific text. CITECHECK grounds each citation in external scholarly sources, compares the retrieved candidate against the citation using a structured LLM verifier, and maps verifier scores to EXACT, MINOR, or MAJOR labels using thresholds tuned on validation data." The authors conclude that "these results suggest that reliable citation hallucination detection benefits from combining scholarly retrieval, structured comparison, and calibrated decision rules, and can serve as a foundation for broader verification pipelines for AI-generated scientific reports."
Risk of bias
Benchmark constructed using only physics citations - may not generalize to other disciplines with different citation practices; Synthetic corruptions generated by GPT-4o-mini may not capture all natural citation hallucination patterns in real LLM-generated content; Retrieval cascade dependent on external service coverage and metadata quality which may exhibit domain-specific or temporal biases; Threshold selection on validation set (n=190) with fixed seed-based partitioning may not generalize to naturally occurring hallucinations; Benchmark is limited to physics domain; may not generalize to other fields with different citation practices. Synthetic corruptions may not capture naturally occurring hallucination patterns. Relies on external metadata sources (CrossRef, Semantic Scholar, OpenAlex, arXiv) which have uneven temporal coverage.; Benchmark is physics-focused and synthetic, may not generalize to other disciplines or naturally occurring LLM-generated errors. Controlled corruptions may not capture full diversity of real-world citation errors. Performance depends heavily on coverage of external scholarly sources, which may have uneven temporal and disciplinary coverage.
Limitations
- "CITECHECK currently depends on the coverage and metadata quality of CrossRef, Semantic Scholar, OpenAlex, arXiv, and web search, so expanding the retrieval cascade to field-specific repositories could improve coverage in domains with specialized publication infrastructure." Additionally, "Our benchmark is also synthetic and focused on physics, which gives controlled labels and clear evaluation but may not capture the full diversity of citation errors in naturally generated reports or in fields with different citation practices." The authors note that "CITECHECK verifies citation existence and metadata fidelity, not whether the cited source supports the surrounding claim."
Open questions raised
- Expanding retrieval cascade to field-specific repositories for domains with specialized publication infrastructure
- Extending benchmark to naturally occurring LLM-generated bibliographies beyond synthetic corruptions
- Testing generalization across additional scientific disciplines beyond physics
- Combining citation verification with downstream claim-citation alignment for end-to-end verification of AI-generated scientific text
- Extending the benchmark to naturally occurring LLM-generated bibliographies and additional disciplines beyond physics; combining CITECHECK with downstream claim–citation alignment for end-to-end verification of AI-generated scientific text; expanding the retrieval cascade to field-specific repositories for improved domain coverage.
- Expanding the retrieval cascade to field-specific repositories for improved coverage in domains with specialized publication infrastructure. Extending the benchmark to naturally occurring LLM-generated bibliographies and additional disciplines beyond physics to test generalization. Combining citation verification with downstream claim-citation alignment toward end-to-end verification of AI-generated scientific text.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations