GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
Zuyao Xu, Yuqi Qiu, Lu Sun, Fasheng Miao, Fubin Wu, Xinyi Wang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Three complementary experiments: (1) LLM Benchmark - systematic evaluation of 13 state-of-the-art LLMs across 40 computer science domains, generating 375,440 citations verified via CITEVERIFIER; (2) Archival Analysis - large-scale audit of 2.2 million citations from 56,381 published papers spanning 2020-2025 across AI/ML and Security venues, with three-stage verification process (automated examination, manual verification by 16 trained assistants, false negative estimation); (3) User Study - survey of 97 researchers across various roles and research domains, with 94 valid responses after quality checks for inconsistent responses.
Sample
N = 56570, 14 groups
Primary method
Descriptive statistics (frequencies, percentages) for survey responses. Empirical cumulative distribution function (ECDF) analysis for similarity scores (Figure 2). Levenshtein distance computation for title similarity matching. The study employed a similarity threshold θ = 0.9 for classification as Valid/Invalid, empirically calibrated using real-paper and LLM-generated citation distributions. Exponential function fitting with R² = 0.94 to model relationship between publication year and hallucination count. Manual verification by 16 trained research assistants with double-checking for flagged citations. Validation sampling: 400 citations from valid pool sampled for quality check (95% confidence, 5% margin of error).
Main result
The study found that "all models hallucinate citations at rates ranging from 14.23% to 94.93%" across 13 LLMs evaluated on citation generation. Additionally, the archival analysis revealed that "1.07% of papers contain invalid citations, with an 80.9% increase in 2025" compared to the 2020-2024 average, and the user study found "87.2% use AI-powered tools in their workflows, 76.7% of reviewers do not thoroughly check references, and 74.5% view peer review as ineffective at catching citation errors."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Empiricist - quantitative measurement of citation validity through automated verification, benchmarking, and survey methods
Author conclusions
The authors conclude: "We presented the first comprehensive investigation of citation validity in the age of LLMs, mapping how ghost citations flow from generation through adoption, review failure, and propagation. We develop CITEVERIFIER as an open-source tool that establishes a measurement baseline and clarifies the systemic nature of the threat. We hope our work motivates the community to treat citation integrity as a critical priority and call for immediate efforts to safeguard the scholarly record before ghost citations become an intractable problem." They further argue that "ghost citations represent a systemic threat to academic integrity, and call for coordinated efforts from community to address this challenge."
Risk of bias
Selection bias in user study recruitment (social media + targeted email sampling from 300 researchers); Social desirability bias in self-reported survey responses (91.5% report responsibility attribution vs. actual behavior); False negative bias in citation verification due to title-only matching approach; False positive bias from OCR errors and poorly indexed legitimate papers; Detection limitations from title-similarity-only verification approach; Social-desirability bias in user survey responses (self-reported verification practices may overstate diligence); Selection bias in survey recruitment (researchers from top-tier venues may not represent broader research community); Potential false negatives in automated citation verification (LLM-based reparser could miss truly invalid citations); False positives risk: legitimate but poorly indexed citations may be classified as invalid; Heterogeneity in bibliographic database coverage across disciplines; LLM benchmark does not reflect real-world citation generation context (fixed-batch generation differs from in-situ writing); Manual verification team may have implicit biases in classifying ambiguous citations; Social desirability bias in survey responses regarding verification practices; Selection bias in survey recruitment (public social media and targeted emails to specific venue authors/PC members); Potential false negatives in automated verification due to title-similarity-only matching; Potential false positives where poorly indexed but legitimate papers are classified as invalid; OCR errors and parsing heterogeneity in PDF extraction; Database coverage limitations (no single bibliographic database provides comprehensive coverage)
Limitations
- "Our benchmark prompts models to generate fixed-number citation per domain
- This ensures fair cross-model comparison but differs from real-world use, where citations are produced while drafting arguments
- Our hallucination rates are therefore a controlled baseline, and may not directly reflect real-world prevalence." Additionally, "Our pipeline verifies only title similarity, reflecting our goal of detecting ghost citations (references that cannot be traced or do not exist) rather than citation errors
- This conservative approach may undercount hallucinations that closely resemble real papers, while some flagged citations may be legitimate but poorly indexed works." The authors also note "Survey response biases
- Our survey relies on self-reported behavior, which is susceptible to social-desirability bias: respondents may over-report diligence and under-report risky practices."
Open questions raised
- Scale of the problem was unknown - lack of empirical measurement of how often LLMs fabricate citations
- No scalable way to detect citation validity due to format heterogeneity and need to verify against multiple bibliographic sources
- Lack of understanding of how invalid citations pass through researcher and peer review processes
- Absence of systematic, large-scale analysis of ghost citation prevalence across models and published literature
- Unclear understanding of why verification behaviors fail to prevent citation propagation
- The authors identify three fundamental research gaps: (1) unknown scale of the problem—no empirical measurement of how often LLMs fabricate citations and how many invalid citations have entered published literature; (2) absence of scalable detection methods for citations validity, as citations deviate from standard formats and require verification against multiple bibliographic sources; (3) lack of understanding of how invalid citations pass through researcher draft and peer review processes to reach publication. The paper calls for investment in detection research, regular measurement studies to track trends, and development of shared verification infrastructure and open validation APIs.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations