12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms

JV Roig · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical measurement study using RIKER (Retrieval Intelligence and Knowledge Extraction Rating) methodology.

Sample

N = 4264, 7 groups

Primary method

Descriptive statistics (means, standard deviations, coefficients of variation), cross-hardware comparison using absolute differences in accuracy and fabrication rates, temperature effect analysis using range-based comparisons (Temp Spread = range of overall accuracy across four temperature settings), correlation analysis comparing grounding accuracy vs. fabrication rates, monotonic trend analysis for coherence loss by temperature. No inferential statistical tests (e.g., ANOVA, t-tests) or hypothesis testing reported.

Main result

The study found that "even under optimal conditions, every model fabricates answers at a non-trivial rate, and fabrication rises steeply with context length." At 32K context, "only 7 of 35 models (20%) achieve a fabrication rate below 10%, and only 2 models (GLM 4.5 at 1.19% and GLM 4.5 Air at 3.37%) stay below 5%." Additionally, "the conventional wisdom of setting temperature to zero for factuality in enterprise document Q&A scenarios is not universally supported—while T=0.0 yields the best overall accuracy in a majority of cases, it can increase both fabrication rates and coherence loss."

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

"This paper set out to answer a deceptively simple question: how much do LLMs hallucinate when answering questions about documents they have been given? Evaluating 35 open-weight models across three context lengths (32K, 128K, 200K), four temperatures, and three hardware platforms—consuming 172 billion tokens across more than 4,000 runs—we find that the answer is 'substantially, and unavoidably.' Even under optimal conditions—best model, best temperature, temperature chosen specifically to minimize fabrication—the floor is non-zero and rises steeply with context length. At 32K, the best model (GLM 4.5) fabricates 1.19% of answers, top-tier models fabricate 5-7%, and the median model fabricates roughly 25%. At 128K, the floor nearly triples to 3.19% and only 5 of 26 tested models remain below 10% fabrication. At 200K, no model stays below 10%."

Risk of bias

Single evaluation framework (RIKER only) - not independently replicated; Language bias - English only evaluation; Model selection bias - open-weight models only, excludes proprietary models (GPT-4, Claude, Gemini); Task specificity bias - document Q&A only, may not generalize to other LLM applications; Temperature sampling bias - coarse grain sampling (0.0, 0.4, 0.7, 1.0) may miss optimal intermediate values; Hardware platform variation - DeepSeek V3.1 anomaly at 32K on MI300X attributed to vLLM version differences; Single evaluation framework (RIKER only) - findings not independently replicated; Model selection bias - only open-weight models tested via vLLM; Application-specific bias - document Q&A only, may not generalize to other tasks; Temperature sampling bias - coarse granularity (0.0, 0.4, 0.7, 1.0) may miss important intermediate effects; Context length bias - evaluation capped at 200K tokens, many models advertise 1M+ support; vLLM serving artifact - DeepSeek V3.1 at 32K showed anomalous cross-platform divergence likely due to vLLM version differences; Hardware-specific bias - Gaudi3 tested on fewer models than H200 and MI300X; Single evaluation framework (RIKER) - results not independently validated with alternative frameworks; Model selection bias - only open-weight models tested, no proprietary models (GPT-4, Claude, Gemini); Language bias - English-only evaluation; results may not generalize to other languages; Application scope bias - document Q&A specific; findings may not transfer to other LLM applications; Temperature sampling bias - coarse granularity (0.0, 0.4, 0.7, 1.0) may miss optimal intermediate temperatures; Hardware consistency artifact - DeepSeek V3.1 at 32K showed anomalous MI300X results potentially due to vLLM version differences

Limitations

  • The authors state: "Single evaluation framework
  • All results are based on the RIKER methodology
  • While RIKER's ground-truth-first approach avoids many pitfalls of traditional evaluation (see Section 3), our findings have not been independently replicated with alternative frameworks
  • English only
  • All generated documents and questions are in English
  • Fabrication rates and context degradation patterns may differ for other languages, particularly for models with varying multilingual training data proportions

Open questions raised

  • Finer temperature granularity - identify optimal temperature sweet spots in the 0.0-0.4 range
  • Longer context evaluation - extend RIKER evaluations to 400K and beyond tokens
  • Repetition penalty as alternative to temperature - test whether repetition penalty parameters prevent coherence loss without accuracy trade-offs
  • Multilingual evaluation - extend RIKER methodology to non-English languages
  • Fabrication-specific training interventions - investigate what training approaches produce low fabrication rates in GLM and MiniMax families
  • Finer temperature granularity in the 0.0-0.4 range to identify optimal temperature sweet spots
Extracted from: pdfAgreement 61%

Explore related topics

Related papers