Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models
Aryan Khurana, Aravind Ramana, Dhruv Kumar · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale multi-domain benchmarking study using a fully balanced 2×2 factorial design crossing claim veracity (true vs.
Sample
N = 220564, 8 groups
Primary method
Primary metric: hallucination rate per condition (proportion of hallucinated outputs over non-refused responses). Effect characterization using lift (absolute pp difference) and Cohen's d with conventional thresholds (d=0.2/0.5/0.8 for small/medium/large effects). Confidence intervals via normal approximation to binomial. Inter-rater reliability: Cohen's kappa. P-values reported for significance testing but effect sizes (Cohen's d) foregrounded given large sample sizes. Refusal rates (below 2%) excluded from denominator. Stratified sampling for 15K subset to preserve 2×2 balance, domain, template, prestige tier, and author demographic distributions.
Main result
The study found that "citation presence, fabricated or real, increases hallucination across all seven models" tested. Most critically, "the effect peaks when fabricated citations accompany true claims, reaching 35–77% hallucination in general knowledge" domain, with lifts of "+3 to 22 percentage points" over true-claim baselines. The authors note that "every model is more likely to deny a correct fact than in any other experimental condition" when fabricated citations accompany true claims.
Reports effect sizes and confidence intervals.
Research paradigm
Empirical quantitative (experimental benchmarking)
Author conclusions
The authors conclude: "Across all seven models tested, adding a citation—fabricated or real—increases hallucination above the no-citation baseline. The effect is most consequential in a condition prior work has not examined: fabricated citations paired with true claims. In this condition every model is more likely to deny a correct fact than in any other experimental condition, with lifts of +3.23 to +22.29 pp over true-claim baselines and near-ceiling hallucination rates (35–77%) in the general knowledge domain. This is not a failure of factual knowledge but of epistemic reasoning under authority pressure." They further note that "susceptibility on true claims does not track model size or capability" and that "mitigation will require models that treat a citation as evidence rather than authority."
Risk of bias
Judge model limitation: Qwen3-8B cannot independently verify cited sources without retrieval access; General knowledge domain: Back-filled citation metadata from other domains (not actual Wikipedia citations); Sample size differences: Full-dataset evaluation limited to smaller open-weight models (3B-8B); API models evaluated only on 15K subset; Computational constraints: Single GPU evaluation limits generalizability; Author demographics proxy: Country-coded surnames are coarse proxy lacking granularity; Template coverage gap: 40 templates do not cover web-native or informal citation formats; Judge model reliability: Validation only on 1,500 samples; Cohen's kappa of 0.83 indicates room for disagreement; Judge model reliability: LLM-based evaluation without retrieval access cannot independently verify citations; Metadata backfilling for general knowledge domain introduces artificial citation-claim mismatches; Sample size variation: Three models evaluated on full dataset (~220K prompts), four on balanced 15K subset; Non-linear model scaling effects across different model families and sizes; Demographic proxy limitation: Country-coded surnames do not capture fine-grained identity signals; Subset representativeness: 15K subset uses proportional stratified sampling but may not perfectly replicate full-dataset distributions; Judge model bias: Qwen3-8B cannot independently verify cited sources or retrieve external information; Back-filled metadata: General knowledge citations use fabricated metadata from other domains rather than real sources; Computational constraints: Subset evaluation of proprietary models limits direct comparability; Template bias: 40 templates do not cover web-native or informal citation formats; Demographic proxy limitations: Country-coded surnames are coarse proxies for perceived demographic identity
Limitations
- The authors explicitly state multiple limitations: "All model outputs are evaluated using Qwen3-8B
- As with all LLM-based judges without retrieval access, it cannot verify whether a cited source exists or supports the attributed claim." Additionally, "True citations for FEVER-sourced claims use author, venue, and year metadata back-filled from other domain citation pools, since Wikipedia articles lack structured academic citation records
- These entries are flagged as citation matches claim = False, and results for this condition should be interpreted accordingly." The study was also "conducted on a single university-owned GPU with 50GB VRAM, constraining full-dataset evaluation to open-source models in the 3B–8B range." Furthermore, "Country-coded surnames are a coarse proxy for perceived demographic identity and do not capture finer-grained signals such as name familiarity or intersectional combinations." Finally, "The 40 templates do not cover web-native or informal citation formats such as hyperlinks, social media references, or conversational citations."
Open questions raised
- Mechanistic analysis via hidden state or attention inspection to clarify why citation signals disrupt true-claim processing and the suppression/amplification split across model families
- Development and evaluation of prompt-based, fine-tuning-based, or architectural mitigations against citation deference
- Expanded demographic analysis at finer granularity beyond country-coded surnames
- Evaluation of larger frontier models beyond those tested
- Retrieval-augmented judge with access to live citation database for more reliable source verification
- Mechanistic analysis via hidden state or attention inspection to clarify why citation signals disrupt true-claim processing and why suppression/amplification split emerges across model families
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations