12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

Khashayar Khajavi, Shaghayegh Sadeghi, Rise Adhikari, Alexander Tessier · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Hybrid computational framework combining information retrieval, structured LLM-based verification, and threshold-based classification.

Main result

On the held-out test set, CITECHECK achieves 88.7 macro-F1 and 88.9% accuracy, outperforming GPT, Claude, and Gemini baselines, including web-search and few-shot variants. "CITECHECK improves over the strongest baseline by 5.8 macro-F1 points and 5.7 accuracy points." The results demonstrate that "reliable citation verification benefits from combining scholarly retrieval, structured LLM-based comparison, and calibrated decision rules."

Research paradigm

computational validation and empirical evaluation

Author conclusions

"We introduced CITECHECK, a framework for detecting hallucinated citations in scientific text. CITECHECK grounds each citation in external scholarly sources, compares the retrieved candidate against the citation using a structured LLM verifier, and maps verifier scores to EXACT, MINOR, or MAJOR labels using thresholds tuned on validation data." The authors conclude that "these results suggest that reliable citation hallucination detection benefits from combining scholarly retrieval, structured comparison, and calibrated decision rules, and can serve as a foundation for broader verification pipelines for AI-generated scientific reports."

Risk of bias

Benchmark constructed using only physics citations - may not generalize to other disciplines with different citation practices; Synthetic corruptions generated by GPT-4o-mini may not capture all natural citation hallucination patterns in real LLM-generated content; Retrieval cascade dependent on external service coverage and metadata quality which may exhibit domain-specific or temporal biases; Threshold selection on validation set (n=190) with fixed seed-based partitioning may not generalize to naturally occurring hallucinations; Benchmark is limited to physics domain; may not generalize to other fields with different citation practices. Synthetic corruptions may not capture naturally occurring hallucination patterns. Relies on external metadata sources (CrossRef, Semantic Scholar, OpenAlex, arXiv) which have uneven temporal coverage.; Benchmark is physics-focused and synthetic, may not generalize to other disciplines or naturally occurring LLM-generated errors. Controlled corruptions may not capture full diversity of real-world citation errors. Performance depends heavily on coverage of external scholarly sources, which may have uneven temporal and disciplinary coverage.

Limitations

  • "CITECHECK currently depends on the coverage and metadata quality of CrossRef, Semantic Scholar, OpenAlex, arXiv, and web search, so expanding the retrieval cascade to field-specific repositories could improve coverage in domains with specialized publication infrastructure." Additionally, "Our benchmark is also synthetic and focused on physics, which gives controlled labels and clear evaluation but may not capture the full diversity of citation errors in naturally generated reports or in fields with different citation practices." The authors note that "CITECHECK verifies citation existence and metadata fidelity, not whether the cited source supports the surrounding claim."

Open questions raised

  • Expanding retrieval cascade to field-specific repositories for domains with specialized publication infrastructure
  • Extending benchmark to naturally occurring LLM-generated bibliographies beyond synthetic corruptions
  • Testing generalization across additional scientific disciplines beyond physics
  • Combining citation verification with downstream claim-citation alignment for end-to-end verification of AI-generated scientific text
  • Extending the benchmark to naturally occurring LLM-generated bibliographies and additional disciplines beyond physics; combining CITECHECK with downstream claim–citation alignment for end-to-end verification of AI-generated scientific text; expanding the retrieval cascade to field-specific repositories for improved domain coverage.
  • Expanding the retrieval cascade to field-specific repositories for improved coverage in domains with specialized publication infrastructure. Extending the benchmark to naturally occurring LLM-generated bibliographies and additional disciplines beyond physics to test generalization. Combining citation verification with downstream claim-citation alignment toward end-to-end verification of AI-generated scientific text.
Data: 982-citation physics benchmark with controlled corruptions (stated as available upon request: "Code and data are available upon request."); "Code and data are available upon request." A 982-citation physics benchmark with controlled corruptions is constructed but conditional availability is noted.; Code and data available upon request (as stated in footnote 1 of paper). The 982-citation physics benchmark with controlled corruptions is mentioned as constructed by the authors.Code: Code availability stated as "upon request" - no specific GitHub or repository URL provided; Code available upon request (not publicly deposited in GitHub or similar repository at time of publication); Code and data are available upon request (repository location not explicitly provided in paper)Extracted from: pdfAgreement 38%

Explore related topics

Related papers