12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

Rohan N. Pradhan, Steve Goley · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Factorial behavioral experiment across five language model families (Claude, Qwen, OLMo) and three professional domains (venture capital, marketing, public health).

Sample

N = 1000000, 9 groups

Primary method

Source Preference Index (SPI) analysis: (ŷ - c)/(a - c) where ŷ is model estimate, a is focal value, c is consensus mean; Correct Identification Rate (CIR): fraction of reviews identifying specific flaw vs. generic concerns; Reasoning trace classification: LLM-as-judge with Cohen's κ agreement validation (κ=0.74-0.90 for primary labels); Linear difference-of-means probes on residual-stream activations with cross-domain transfer evaluation (6-fold train/test split); ROME-style causal tracing: corrupt presentation embeddings with calibrated Gaussian noise, restore layer-by-layer, measure recovery ρ; Direct logit attribution (DLA): project component outputs onto analyst-vs-consensus logit direction; Non-linear MLP probes (2-layer, d_model → 256 → 1, ReLU dropout); Bootstrap confidence intervals (5th-95th percentile winsorization for restore curves); Jargon-density control: 1-D logistic regression on density, residualization via linear regression; Permutation testing (1,000 label shuffles for probe significance)

Main result

The study found that "across all social configurations, adding valid methodology with statistics to the focal source's claim shifts it from the weakest source to the dominant one; adding impossible statistics produces nearly the same shift." Models treat fabricated statistics as if they were valid during multi-source synthesis, despite detecting them reliably in isolation. "When the focal source is socially isolated, adding a plausible methodology description shifts the model most of the way from consensus toward the focal source's estimate. Substituting an impossible statistic for the valid one produces close to the same effect: a claim backed by an impossibly narrow CI exerts the same pull as one backed by a valid CI."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical quantitative (experimental and mechanistic analysis)

Author conclusions

"As language models increasingly serve as epistemic proxies, the question of whether they evaluate the quality of their sources becomes urgent. We use epistemic alignment to name this gap: between the evaluation the model performs automatically (methodology register) and the evaluation it does not perform at all (numeric validity). Closing this gap may require training on data where premises are sometimes invalid and need to be verified, a form of supervision that current pipelines may not provide."

Risk of bias

Synthetic experimental setting may not reflect real-world source evaluation contexts; Limited architectural diversity (primarily Transformer-based models); Partial mechanistic replication (OLMo 3.1 in addition to Qwen 32B); Internal validity: LLM-as-judge classification may introduce systematic bias; author validation on 100 samples shows κ=0.74-0.90 for key labels but κ=0.39 for defers_to label, suggesting threshold-dependent bias; Construct validity: Source Preference Index normalized by focal-consensus gap may not capture anchoring independent of genuine credibility judgment; Ecological validity: Fixed four-source structure and synthetic domain scenarios limit generalizability to real-world multi-source synthesis; Mechanistic inference: Linear probes cannot rule out indirect compensatory pathways; causal tracing on OLMo uses out-of-distribution prefill (<think></think><estimate>) that may attenuate signals; Confounding: Token-length differences between specious and plausible conditions partially controlled via residualization but not fully eliminated; Model selection: Five models tested but mechanistic analysis primarily on Qwen 32B with partial OLMo 3.1 replication; Experimental context may not reflect real-world deployment conditions (fixed four-source structure); Mechanistic analysis conducted primarily on one architecture (Qwen 32B) with partial replication; Synthetic domain scenarios rather than real organizational data; Model selection bias: only five model families tested

Limitations

  • "Our design trades ecological breadth for experimental control: the fixed four-source structure enables a full factorial but leaves open whether the same vulnerability emerges in less structured multi-source settings
  • The mechanistic analysis, conducted on Qwen 32B with partial replication on OLMo 3.1, characterizes direct logit contributions and cannot rule out indirect pathways that partially compensate."

Open questions raised

  • Mitigation strategies for the fabrication vulnerability remain undeveloped
  • Generalization to less structured multi-source settings unknown
  • Indirect computational pathways not fully characterized
  • Training approaches that develop numeric verification capability not identified
  • Whether the vulnerability emerges in less structured multi-source settings beyond the fixed four-source paradigm
  • Indirect computational pathways that may partially compensate for the identified consensus-gated mechanism
Data: Datasets not made available as explicit downloadable artifacts; all 2,304 conditions per domain reproducible from factorial specification and verbatim templates provided in appendicesCode: Not mentionedExtracted from: pdfAgreement 62%

Explore related topics

Related papers