12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI Hallucinations in Retrieval-Augmented and Generative Systems: A Rigorous Review of Definitions, Failure Mechanisms, Evaluation, and Mitigation Strategies

Toufik Mzili, Ilyass Mzili, Ahmed Abatal, Zahra Oughannou, Adarsh Kumar Arya, Mohamed Kurdi · EDRAAK · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.70470/edraak/2026/004

Methodology & findings

Study design

Structured narrative-review approach using 70 references from a supplied bibliography spanning journal articles, conference papers, preprints, essays, and book chapters.

Sample

N = 70, 4 groups

Main result

The study found that "hallucination is best treated as a family of failures rather than a single phenomenon, since false information, grounding conflicts, reference errors, and multimodal distortions do not arise or behave in identical ways." The research demonstrates that hallucinations persist across three levels of analysis: the model level (probabilistic next-token generation), the pipeline system level (retrieval, ranking, chunking, prompt construction), and the human-governance level (trust, anthropomorphism, responsibility attribution). The authors conclude that "RAG operates as a solution to hallucinations in AI yet the analyzed research demonstrates it operates as a system which moves failure points instead of completely eliminating them."

Reports effect sizes.

Research paradigm

Critical interpretivism with systems analysis

Author conclusions

The authors conclude that "The research materials reviewed in this section establish that hallucination problem as a major problem which continues to exist after retrieval systems have been implemented." They emphasize that "The system requires disciplined design which includes curated data and strong retrieval capabilities and explicit evidence display and constrained generation and uncertainty quantification and post-hoc verification and accountable human oversight." Most critically, they state: "In short, the field needs fewer vague promises and more rigorous measurement. Fluency is cheap. Grounded reliability is the real work."

Risk of bias

Heterogeneous evidence quality across sources (experimental data, structured evaluations, reflective essays, position papers); Lack of consistent empirical information throughout corpus; Sources treated inequivalently (technical surveys and empirical investigations given more weight than commentary pieces); Publication bias toward recent literature (sharp increase after 2023); Multidisciplinary corpus may introduce domain-specific reporting biases; Benchmark fragmentation across studies using different error definitions; Source heterogeneity bias: The corpus mixes experimental data, structured evaluations, reflective essays, position papers, and conceptual reviews, treated unequally; Publication bias: Sharp increase in publications after 2023 may reflect research trends rather than actual problem prevalence; Disciplinary bias: Multidisciplinary corpus may privilege technical and computer science perspectives over applied domain perspectives; Definition bias: Inconsistent definitions of 'hallucination' across sources creates measurement inconsistency; Selection bias in corpus: Authors selected 70 references from a Word file; criteria for inclusion not fully specified; Publication bias: corpus reflects available literature which may skew toward positive results or recent high-profile cases; Source heterogeneity: mixing of empirical studies with commentary and position papers may introduce interpretation bias; Selection bias: the 70 references were provided, not systematically searched, potentially missing relevant literature; Thematic concentration bias: literature clustered around definition, evaluation, mitigation, and governance topics may exclude other relevant perspectives; Domain representation bias: multidisciplinary sources may not represent equal empirical rigor across domains

Limitations

  • The authors state that "The main restriction needs to be directly stated
  • The corpus contains extensive content but it lacks consistent empirical information throughout its material
  • The sources include experimental data from some sources and others present structured evaluations but other sources contain reflective essays and position papers and conceptual critical reviews." Additionally, they note that "because the input bibliography spans journal papers, conference papers, preprints, essays, book chapters, and commentary pieces, a formal meta-analysis would not be methodologically appropriate." The field shows "inconsistent methodological approaches despite its fast-growing number of published works" with "benchmark fragmentation" persisting because "Multiple research papers employ various methods to determine error quantities because they use different error definitions."

Open questions raised

  • Benchmark fragmentation: Multiple studies use different error definitions and evaluation methods, making comparison of results difficult
  • Model-centric focus: Most studies treat hallucination as a generator issue when actual errors stem from data processing, retrieval performance, evidence selection, interface design, and user comprehension
  • Insufficient longitudinal research: Most existing work fails to meet quality standards for studying how users adapt to hallucinations over time
  • Regulatory gaps outpacing technical clarity: Development of formal standards for error tolerance, domain-specific abstention, and evidence presentation remains insufficient
  • Need for comprehensive testing: Future research requires benchmark suites spanning multiple layers, domain-weighted harm metrics, mitigation stack comparisons, and testing with actual user scenarios rather than benchmark prompts
  • Need for domain-specific guidelines: Particular rules must be developed determining when creative works should include fiction versus when vital content requires validation
Extracted from: pdfAgreement 70%

Explore related topics

Related papers