12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges

Big Data and Cognitive Computing · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
I
Evidence
13
Citations
21.26
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/bdcc9120320

Methodology & findings

Study design

Systematic literature review conducted according to PRISMA 2020 guidelines.

Primary method

Systematic literature review following PRISMA 2020

Main result

The review synthesized 128 studies on retrieval-augmented generation (RAG) and found that "methods have shifted from DPR+seq2seq baselines to modular, policy-driven RAG with hybrid/structure-aware retrieval, uncertainty-triggered loops, memory, and emerging multimodality." Additionally, "evaluation remains overlap-heavy (EM/F1), with increasing use of retrieval diagnostics (e.g., Recall@k, MRR@k), human judgements, and LLM-as-judge protocols." The evidence base shows that "nearly half of all studies concentrating on knowledge-intensive and open-domain QA settings," with emerging applications in software engineering (10.2%) and medical domains (8.6%).

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "Evidence supports a shift to modular, policy-driven RAG, combining hybrid/structure-aware retrieval, uncertainty-aware control, memory, and multimodality, to improve grounding and efficiency." They further recommend: "(i) holistic benchmarks pairing quality with cost/latency and safety, (ii) budget-aware retrieval/tool-use policies, and (iii) provenance-aware pipelines that expose uncertainty and deliver traceable evidence." They note that "To advance from prototypes to dependable systems" these recommendations are essential for the field's maturation.

Risk of bias

Citation-lag bias from applying lower citation thresholds to 2025 publications; Time-window bias from 2020-2025 restriction favouring earlier, highly cited papers; Language bias (English-only inclusion); Database coverage bias (five libraries only); Domain concentration bias (nearly 50% of studies in QA/knowledge-intensive tasks); Benchmark-centric bias in dataset selection (NQ, HotPotQA, TriviaQA, MS MARCO, Wikipedia); Selection bias from full-text availability requirement; Selection bias toward highly-cited works; Benchmark-centric bias: Literature appears optimized for benchmark achievement rather than real-world deployment metrics; Selection bias in data extraction: Single reviewer performed initial data extraction using RAG framework, though human verification was performed; RAG framework poses risks of hallucination and missing key data

Limitations

  • The authors acknowledge several key limitations: "We note the evidence base may be affected by citation-lag from the inclusion thresholds and by English-only, five-library coverage." Additionally, "Restricting the review to 2020–2025 could introduce a time-window bias by favouring earlier, highly cited papers." The review also notes that "coverage is still uneven, which limits how confidently methods can be lifted from benchmark QA and dropped into domain-specific pipelines without additional tuning," particularly in areas with scattered studies (finance, education, security, biomedical) where "evaluation practices are heterogeneous and often under-specified, making cross-paper comparisons fragile." The authors further state that "many reported gains emphasise retrieval–generation mechanics (e.g., retriever choice, reranking, context shaping) over deployment concerns," and "evidence is thinner for low-resource or safety-critical contexts where privacy, governance, or drift matter as much as accuracy."

Open questions raised

  • Inconsistent evaluation practices across domains limiting cross-paper comparisons
  • Thin evidence base for safety-critical, low-resource, and privacy-governed contexts
  • Lack of standardised reporting of system-level metrics (cost per query, end-to-end latency)
  • Uneven coverage in applied domains (finance, education, security, biomedical)
  • Need for joint reporting of accuracy, cost, robustness, and security in single studies
  • Missing data on parameter-efficient fine-tuning effects on RAG outcomes
Data: Structured database available via Zenodo (DOI: 10.5281/zenodo.17339384; uploaded 13 October 2025); Natural Questions (NQ); HotPotQA; TriviaQA; MS MARCO; Wikipedia; FinanceBenchCode: Python 3.12 scripts for deduplication and export (stated to be made available during peer review and prior to publication)Extracted from: pdf

Explore related topics

Related papers