How Much Do LLMs Hallucinate in Document Q&A Scenarios? A 172-Billion-Token Study Across Temperatures, Context Lengths, and Hardware Platforms
JV Roig · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical measurement study using RIKER (Retrieval Intelligence and Knowledge Extraction Rating) methodology.
Sample
N = 4264, 7 groups
Primary method
Descriptive statistics (means, standard deviations, coefficients of variation), cross-hardware comparison using absolute differences in accuracy and fabrication rates, temperature effect analysis using range-based comparisons (Temp Spread = range of overall accuracy across four temperature settings), correlation analysis comparing grounding accuracy vs. fabrication rates, monotonic trend analysis for coherence loss by temperature. No inferential statistical tests (e.g., ANOVA, t-tests) or hypothesis testing reported.
Main result
The study found that "even under optimal conditions, every model fabricates answers at a non-trivial rate, and fabrication rises steeply with context length." At 32K context, "only 7 of 35 models (20%) achieve a fabrication rate below 10%, and only 2 models (GLM 4.5 at 1.19% and GLM 4.5 Air at 3.37%) stay below 5%." Additionally, "the conventional wisdom of setting temperature to zero for factuality in enterprise document Q&A scenarios is not universally supported—while T=0.0 yields the best overall accuracy in a majority of cases, it can increase both fabrication rates and coherence loss."
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist
Author conclusions
"This paper set out to answer a deceptively simple question: how much do LLMs hallucinate when answering questions about documents they have been given? Evaluating 35 open-weight models across three context lengths (32K, 128K, 200K), four temperatures, and three hardware platforms—consuming 172 billion tokens across more than 4,000 runs—we find that the answer is 'substantially, and unavoidably.' Even under optimal conditions—best model, best temperature, temperature chosen specifically to minimize fabrication—the floor is non-zero and rises steeply with context length. At 32K, the best model (GLM 4.5) fabricates 1.19% of answers, top-tier models fabricate 5-7%, and the median model fabricates roughly 25%. At 128K, the floor nearly triples to 3.19% and only 5 of 26 tested models remain below 10% fabrication. At 200K, no model stays below 10%."
Risk of bias
Single evaluation framework (RIKER only) - not independently replicated; Language bias - English only evaluation; Model selection bias - open-weight models only, excludes proprietary models (GPT-4, Claude, Gemini); Task specificity bias - document Q&A only, may not generalize to other LLM applications; Temperature sampling bias - coarse grain sampling (0.0, 0.4, 0.7, 1.0) may miss optimal intermediate values; Hardware platform variation - DeepSeek V3.1 anomaly at 32K on MI300X attributed to vLLM version differences; Single evaluation framework (RIKER only) - findings not independently replicated; Model selection bias - only open-weight models tested via vLLM; Application-specific bias - document Q&A only, may not generalize to other tasks; Temperature sampling bias - coarse granularity (0.0, 0.4, 0.7, 1.0) may miss important intermediate effects; Context length bias - evaluation capped at 200K tokens, many models advertise 1M+ support; vLLM serving artifact - DeepSeek V3.1 at 32K showed anomalous cross-platform divergence likely due to vLLM version differences; Hardware-specific bias - Gaudi3 tested on fewer models than H200 and MI300X; Single evaluation framework (RIKER) - results not independently validated with alternative frameworks; Model selection bias - only open-weight models tested, no proprietary models (GPT-4, Claude, Gemini); Language bias - English-only evaluation; results may not generalize to other languages; Application scope bias - document Q&A specific; findings may not transfer to other LLM applications; Temperature sampling bias - coarse granularity (0.0, 0.4, 0.7, 1.0) may miss optimal intermediate temperatures; Hardware consistency artifact - DeepSeek V3.1 at 32K showed anomalous MI300X results potentially due to vLLM version differences
Limitations
- The authors state: "Single evaluation framework
- All results are based on the RIKER methodology
- While RIKER's ground-truth-first approach avoids many pitfalls of traditional evaluation (see Section 3), our findings have not been independently replicated with alternative frameworks
- English only
- All generated documents and questions are in English
- Fabrication rates and context degradation patterns may differ for other languages, particularly for models with varying multilingual training data proportions
Open questions raised
- Finer temperature granularity - identify optimal temperature sweet spots in the 0.0-0.4 range
- Longer context evaluation - extend RIKER evaluations to 400K and beyond tokens
- Repetition penalty as alternative to temperature - test whether repetition penalty parameters prevent coherence loss without accuracy trade-offs
- Multilingual evaluation - extend RIKER methodology to non-English languages
- Fabrication-specific training interventions - investigate what training approaches produce low fabrication rates in GLM and MiniMax families
- Finer temperature granularity in the 0.0-0.4 range to identify optimal temperature sweet spots
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations