12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

Aryan Khurana, Aravind Ramana, Dhruv Kumar · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale multi-domain benchmarking study using a fully balanced 2×2 factorial design crossing claim veracity (true vs.

Sample

N = 220564, 8 groups

Primary method

Primary metric: hallucination rate per condition (proportion of hallucinated outputs over non-refused responses). Effect characterization using lift (absolute pp difference) and Cohen's d with conventional thresholds (d=0.2/0.5/0.8 for small/medium/large effects). Confidence intervals via normal approximation to binomial. Inter-rater reliability: Cohen's kappa. P-values reported for significance testing but effect sizes (Cohen's d) foregrounded given large sample sizes. Refusal rates (below 2%) excluded from denominator. Stratified sampling for 15K subset to preserve 2×2 balance, domain, template, prestige tier, and author demographic distributions.

Main result

The study found that "citation presence, fabricated or real, increases hallucination across all seven models" tested. Most critically, "the effect peaks when fabricated citations accompany true claims, reaching 35–77% hallucination in general knowledge" domain, with lifts of "+3 to 22 percentage points" over true-claim baselines. The authors note that "every model is more likely to deny a correct fact than in any other experimental condition" when fabricated citations accompany true claims.

Reports effect sizes and confidence intervals.

Research paradigm

Empirical quantitative (experimental benchmarking)

Author conclusions

The authors conclude: "Across all seven models tested, adding a citation—fabricated or real—increases hallucination above the no-citation baseline. The effect is most consequential in a condition prior work has not examined: fabricated citations paired with true claims. In this condition every model is more likely to deny a correct fact than in any other experimental condition, with lifts of +3.23 to +22.29 pp over true-claim baselines and near-ceiling hallucination rates (35–77%) in the general knowledge domain. This is not a failure of factual knowledge but of epistemic reasoning under authority pressure." They further note that "susceptibility on true claims does not track model size or capability" and that "mitigation will require models that treat a citation as evidence rather than authority."

Risk of bias

Judge model limitation: Qwen3-8B cannot independently verify cited sources without retrieval access; General knowledge domain: Back-filled citation metadata from other domains (not actual Wikipedia citations); Sample size differences: Full-dataset evaluation limited to smaller open-weight models (3B-8B); API models evaluated only on 15K subset; Computational constraints: Single GPU evaluation limits generalizability; Author demographics proxy: Country-coded surnames are coarse proxy lacking granularity; Template coverage gap: 40 templates do not cover web-native or informal citation formats; Judge model reliability: Validation only on 1,500 samples; Cohen's kappa of 0.83 indicates room for disagreement; Judge model reliability: LLM-based evaluation without retrieval access cannot independently verify citations; Metadata backfilling for general knowledge domain introduces artificial citation-claim mismatches; Sample size variation: Three models evaluated on full dataset (~220K prompts), four on balanced 15K subset; Non-linear model scaling effects across different model families and sizes; Demographic proxy limitation: Country-coded surnames do not capture fine-grained identity signals; Subset representativeness: 15K subset uses proportional stratified sampling but may not perfectly replicate full-dataset distributions; Judge model bias: Qwen3-8B cannot independently verify cited sources or retrieve external information; Back-filled metadata: General knowledge citations use fabricated metadata from other domains rather than real sources; Computational constraints: Subset evaluation of proprietary models limits direct comparability; Template bias: 40 templates do not cover web-native or informal citation formats; Demographic proxy limitations: Country-coded surnames are coarse proxies for perceived demographic identity

Limitations

  • The authors explicitly state multiple limitations: "All model outputs are evaluated using Qwen3-8B
  • As with all LLM-based judges without retrieval access, it cannot verify whether a cited source exists or supports the attributed claim." Additionally, "True citations for FEVER-sourced claims use author, venue, and year metadata back-filled from other domain citation pools, since Wikipedia articles lack structured academic citation records
  • These entries are flagged as citation matches claim = False, and results for this condition should be interpreted accordingly." The study was also "conducted on a single university-owned GPU with 50GB VRAM, constraining full-dataset evaluation to open-source models in the 3B–8B range." Furthermore, "Country-coded surnames are a coarse proxy for perceived demographic identity and do not capture finer-grained signals such as name familiarity or intersectional combinations." Finally, "The 40 templates do not cover web-native or informal citation formats such as hyperlinks, social media references, or conversational citations."

Open questions raised

  • Mechanistic analysis via hidden state or attention inspection to clarify why citation signals disrupt true-claim processing and the suppression/amplification split across model families
  • Development and evaluation of prompt-based, fine-tuning-based, or architectural mitigations against citation deference
  • Expanded demographic analysis at finer granularity beyond country-coded surnames
  • Evaluation of larger frontier models beyond those tested
  • Retrieval-augmented judge with access to live citation database for more reliable source verification
  • Mechanistic analysis via hidden state or attention inspection to clarify why citation signals disrupt true-claim processing and why suppression/amplification split emerges across model families
Data: AuthorityBench (220,564 prompts): https://github.com/floatingreeds/AuthorityBench; FEVER (Thorne et al., 2018): General knowledge claims from Wikipedia; SciQ (Welbl et al., 2017): Science exam dataset; CaseHOLD (Zheng et al., 2021): Legal MCQ dataset; MedMCQA (Pal et al., 2022): Medical MCQ dataset; MTEB SciFact (Wadden et al., 2020): Science citations; PubMedQA (Jin et al., 2019): Medical citations; AuthorityBench; FEVER; SciQ; CaseHOLD; MedMCQA; MTEB SciFact; PubMedQA; AuthorityBench (220,564 prompts across four domains) - https://github.com/floating-reeds/AuthorityBench; FEVER (Thorne et al., 2018) - general knowledge claims; SciQ (Welbl et al., 2017) - science exam questions; CaseHOLD (Zheng et al., 2021) - legal claims; MedMCQA (Pal et al., 2022) - medical MCQ dataset; MTEB SciFact (Wadden et al., 2020) - science citations; PubMedQA (Jin et al., 2019) - medical citations with PubMed metadataCode: https://github.com/floatingreeds/AuthorityBench; https://github.com/floating-reeds/AuthorityBench (variant URL cited in text); GitHub: https://github.com/floating-reeds/AuthorityBench; https://github.com/floating-reeds/AuthorityBenchExtracted from: pdfAgreement 50%

Explore related topics

Related papers