12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes

Chen Zhao, Yuan Tang, Yitian Qian · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical measurement study across four LLMs (two proprietary: Claude Sonnet, GPT-4o; two open-weight: Qwen 2.5-14B, LLaMA 3.1-8B) evaluated under five prompting conditions (Baseline, Temporal, Survey, Non-Disclosure, Combo) using deterministic decoding.

Sample

N = 17443, 4 groups

Primary method

Cluster-bootstrap resampling (1,000 resamples over 144 claims) to compute 95% confidence intervals accounting for within-claim correlation. Bootstrap CI of the difference (Δ) in existence rate for pairwise comparisons. Weighted metadata similarity scoring (Equation 1) with fuzzy string matching. Cohen's kappa (κ) for inter-rater agreement with manual audit (κ=0.63 vs. human labels on 100-citation stratified sample). Precision and recall computed for each label category. Proportional reclassification sensitivity analysis for Unresolved citations.

Main result

The study found that "no model, under any condition, achieves an existence rate above 0.50" and that "temporal constraints reduce citation quality more than any other single condition," with Claude Sonnet falling from 0.381 to 0.119. Additionally, "unresolved outcomes constitute 36-61% of citations across nearly all cells," and "the proprietary-open-weight gap is large and statistically clear (Δ = +0.229, 95% CI [0.191, 0.266])," demonstrating that deployment constraints systematically worsen citation hallucination in qualitatively different ways.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist (quantitative measurement and verification)

Author conclusions

The authors conclude that "across four models and five conditions, deployment constraints consistently worsen citation hallucination-but in qualitatively different ways. Temporal constraints sharply reduce verifiability while maintaining format compliance; Survey prompting widens the proprietary-open-weight gap; Non-Disclosure instructions shift errors from 'obviously wrong' into 'hard to tell'; and models keep citing at high volume even as verifiability erodes." They further state: "any LLM-generated reference list should be treated as a draft requiring independent verification against scholarly databases before inclusion in reviews, technical reports, or tooling pipelines."

Risk of bias

Selection bias in claim dataset: 144 prompts randomly sampled from 240 candidate items, potentially introducing domain representation bias; Database coverage bias: Crossref and Semantic Scholar may not index all venues equally (particularly regional or conference proceedings); False positive fabrication rates: Workshop papers, preprints, and regional-venue articles may be incorrectly labeled as fabricated; Model scale confounding: Cannot fully disentangle whether proprietary-open-weight gap reflects training data size, model scale, or both; Single snapshot evaluation: Results characterize specific model versions at a single time point; generalization uncertain as models evolve; Template and prompt sensitivity: Single prompt template per condition and fixed phrasing per claim may not capture paraphrase robustness; Threshold arbitrariness: Label boundaries (0.85/0.60) and scoring weights are design choices that could shift citations between categories; Selection bias: 144 claims randomly sampled from 240-item candidate pool, not representative of all potential citation scenarios; Model scale confound: Cannot disentangle proprietary model scale from proprietary-openweight distinction; Database coverage bias: Crossref and Semantic Scholar incomplete coverage may cause false positive 'fabricated' labels for preprints, workshop papers, regional venues; Prompt sensitivity: Single prompt template and fixed phrasing per claim may not capture paraphrase robustness; Label boundary choice: 0.85/0.60 thresholds and scoring weights are design choices that could shift citations between categories; Temporal snapshot: Evaluation at single model version; proprietary-openweight gap may change with model updates; Domain representativeness: 24 claims per domain group insufficient for fine-grained subdomain analysis; Selection bias: 144 claims randomly sampled from candidate pool of 240; external validity may be limited; Measurement bias: Automated verification pipeline boundaries (0.85/0.60 thresholds) are design choices that could move citations between categories; Model confounding: Cannot fully disentangle model scale from proprietary-open-weight distinction; Database coverage bias: Crossref and Semantic Scholar do not index all literature (preprints, workshop papers, regional venues); Temporal bias: Evaluation at single snapshot; newer model versions may differ; Prompt sensitivity: Single prompt template per condition; alternative phrasings may yield different results

Limitations

  • The authors state that "our label boundaries (0.85/0.60) and scoring weights are design choices applied uniformly
  • alternative settings could move citations between categories" and "we tested at temperature zero with one prompt template per condition and one fixed phrasing per claim
  • results therefore characterize constraint types rather than paraphrase robustness." External validity concerns include: "our 144 English-language claims may lack power for fine-grained comparisons (e.g., individual subdomains), and findings may not generalize to non-English domains or discipline-specific citation norms." Additionally, "our pipeline relies on Crossref and Semantic Scholar
  • neither indexes everything, so some 'fabricated' labels may be false positives (e.g., preprints, workshop papers, or regional-venue articles absent from both databases)." Finally, "we cannot fully disentangle model scale from the proprietary-open-weight distinction" and "the proprietary-open-weight gap could narrow as open-weight models scale up or incorporate retrieval."

Open questions raised

  • Paraphrase robustness: Systematically varying claim phrasing to disentangle prompt sensitivity from constraint effects
  • Retrieval-augmented comparison: Comparing closed-book generation against retrieval-augmented settings under the same constraint regimes
  • Real-world deployment: Embedding the verification pipeline as a real-time post-generation filter (e.g., IDE plugin or CI check) for manuscript drafts
  • Citation-claim alignment: Whether verified citations actually support the claims made (current pipeline does not check this)
  • Expanded database coverage: Adding DBLP/OpenAlex would likely narrow the Unresolved category
  • Temporal generalization: Tracking whether citation reliability improves as new models and retrieval-augmented architectures emerge
Data: Dataset of 144 claims spanning six academic domains (SE & CS: 24 claims, Natural Sciences: 24, Medicine & Health: 24, Social Sciences: 24, Humanities: 24, Interdisciplinary: 24). Replication package available at https://github.com/Zerichen/Citation-Hallucination; Curated dataset of 144 claims spanning six academic domains (SE & CS: 24; Natural Sciences: 24; Medicine & Health: 24; Social Sciences: 24; Humanities: 24; Interdisciplinary: 24). Available via replication package.; 144-claim dataset with domain labels, temporal windows, and seed anchors available in replication packageCode: https://github.com/Zerichen/Citation-Hallucination (verification pipeline and dataset); https://github.com/Zerichen/Citation-Hallucination (verification pipeline code publicly available); https://github.com/Zerichen/Citation-HallucinationExtracted from: pdfAgreement 53%

Explore related topics

Related papers