Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes
Chen Zhao, Yuan Tang, Yitian Qian · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical measurement study across four LLMs (two proprietary: Claude Sonnet, GPT-4o; two open-weight: Qwen 2.5-14B, LLaMA 3.1-8B) evaluated under five prompting conditions (Baseline, Temporal, Survey, Non-Disclosure, Combo) using deterministic decoding.
Sample
N = 17443, 4 groups
Primary method
Cluster-bootstrap resampling (1,000 resamples over 144 claims) to compute 95% confidence intervals accounting for within-claim correlation. Bootstrap CI of the difference (Δ) in existence rate for pairwise comparisons. Weighted metadata similarity scoring (Equation 1) with fuzzy string matching. Cohen's kappa (κ) for inter-rater agreement with manual audit (κ=0.63 vs. human labels on 100-citation stratified sample). Precision and recall computed for each label category. Proportional reclassification sensitivity analysis for Unresolved citations.
Main result
The study found that "no model, under any condition, achieves an existence rate above 0.50" and that "temporal constraints reduce citation quality more than any other single condition," with Claude Sonnet falling from 0.381 to 0.119. Additionally, "unresolved outcomes constitute 36-61% of citations across nearly all cells," and "the proprietary-open-weight gap is large and statistically clear (Δ = +0.229, 95% CI [0.191, 0.266])," demonstrating that deployment constraints systematically worsen citation hallucination in qualitatively different ways.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist (quantitative measurement and verification)
Author conclusions
The authors conclude that "across four models and five conditions, deployment constraints consistently worsen citation hallucination-but in qualitatively different ways. Temporal constraints sharply reduce verifiability while maintaining format compliance; Survey prompting widens the proprietary-open-weight gap; Non-Disclosure instructions shift errors from 'obviously wrong' into 'hard to tell'; and models keep citing at high volume even as verifiability erodes." They further state: "any LLM-generated reference list should be treated as a draft requiring independent verification against scholarly databases before inclusion in reviews, technical reports, or tooling pipelines."
Risk of bias
Selection bias in claim dataset: 144 prompts randomly sampled from 240 candidate items, potentially introducing domain representation bias; Database coverage bias: Crossref and Semantic Scholar may not index all venues equally (particularly regional or conference proceedings); False positive fabrication rates: Workshop papers, preprints, and regional-venue articles may be incorrectly labeled as fabricated; Model scale confounding: Cannot fully disentangle whether proprietary-open-weight gap reflects training data size, model scale, or both; Single snapshot evaluation: Results characterize specific model versions at a single time point; generalization uncertain as models evolve; Template and prompt sensitivity: Single prompt template per condition and fixed phrasing per claim may not capture paraphrase robustness; Threshold arbitrariness: Label boundaries (0.85/0.60) and scoring weights are design choices that could shift citations between categories; Selection bias: 144 claims randomly sampled from 240-item candidate pool, not representative of all potential citation scenarios; Model scale confound: Cannot disentangle proprietary model scale from proprietary-openweight distinction; Database coverage bias: Crossref and Semantic Scholar incomplete coverage may cause false positive 'fabricated' labels for preprints, workshop papers, regional venues; Prompt sensitivity: Single prompt template and fixed phrasing per claim may not capture paraphrase robustness; Label boundary choice: 0.85/0.60 thresholds and scoring weights are design choices that could shift citations between categories; Temporal snapshot: Evaluation at single model version; proprietary-openweight gap may change with model updates; Domain representativeness: 24 claims per domain group insufficient for fine-grained subdomain analysis; Selection bias: 144 claims randomly sampled from candidate pool of 240; external validity may be limited; Measurement bias: Automated verification pipeline boundaries (0.85/0.60 thresholds) are design choices that could move citations between categories; Model confounding: Cannot fully disentangle model scale from proprietary-open-weight distinction; Database coverage bias: Crossref and Semantic Scholar do not index all literature (preprints, workshop papers, regional venues); Temporal bias: Evaluation at single snapshot; newer model versions may differ; Prompt sensitivity: Single prompt template per condition; alternative phrasings may yield different results
Limitations
- The authors state that "our label boundaries (0.85/0.60) and scoring weights are design choices applied uniformly
- alternative settings could move citations between categories" and "we tested at temperature zero with one prompt template per condition and one fixed phrasing per claim
- results therefore characterize constraint types rather than paraphrase robustness." External validity concerns include: "our 144 English-language claims may lack power for fine-grained comparisons (e.g., individual subdomains), and findings may not generalize to non-English domains or discipline-specific citation norms." Additionally, "our pipeline relies on Crossref and Semantic Scholar
- neither indexes everything, so some 'fabricated' labels may be false positives (e.g., preprints, workshop papers, or regional-venue articles absent from both databases)." Finally, "we cannot fully disentangle model scale from the proprietary-open-weight distinction" and "the proprietary-open-weight gap could narrow as open-weight models scale up or incorporate retrieval."
Open questions raised
- Paraphrase robustness: Systematically varying claim phrasing to disentangle prompt sensitivity from constraint effects
- Retrieval-augmented comparison: Comparing closed-book generation against retrieval-augmented settings under the same constraint regimes
- Real-world deployment: Embedding the verification pipeline as a real-time post-generation filter (e.g., IDE plugin or CI check) for manuscript drafts
- Citation-claim alignment: Whether verified citations actually support the claims made (current pipeline does not check this)
- Expanded database coverage: Adding DBLP/OpenAlex would likely narrow the Unresolved category
- Temporal generalization: Tracking whether citation reliability improves as new models and retrieval-augmented architectures emerge
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations