12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges

Big Data and Cognitive Computing · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
13
Citations
21.26
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/bdcc9120320

Methodology & findings

Study design

Systematic literature review conducted according to PRISMA 2020 guidelines.

Primary method

Systematic literature review following PRISMA 2020

Main result

The review synthesized 128 studies on retrieval-augmented generation (RAG) and found that "methods have shifted from DPR+seq2seq baselines to modular, policy-driven RAG with hybrid/structure-aware retrieval, uncertainty-triggered loops, memory, and emerging multimodality." Additionally, "evaluation remains overlap-heavy (EM/F1), with increasing use of retrieval diagnostics (e.g., Recall@k, MRR@k), human judgements, and LLM-as-judge protocols." The evidence base shows that "nearly half of all studies concentrating on knowledge-intensive and open-domain QA settings," with emerging applications in software engineering (10.2%) and medical domains (8.6%).

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "Evidence supports a shift to modular, policy-driven RAG, combining hybrid/structure-aware retrieval, uncertainty-aware control, memory, and multimodality, to improve grounding and efficiency." They further recommend: "(i) holistic benchmarks pairing quality with cost/latency and safety, (ii) budget-aware retrieval/tool-use policies, and (iii) provenance-aware pipelines that expose uncertainty and deliver traceable evidence." They note that "To advance from prototypes to dependable systems" these recommendations are essential for the field's maturation.

Risk of bias

Citation-lag bias from applying lower citation thresholds to 2025 publications; Time-window bias from 2020-2025 restriction favouring earlier, highly cited papers; Language bias (English-only inclusion); Database coverage bias (five libraries only); Domain concentration bias (nearly 50% of studies in QA/knowledge-intensive tasks); Benchmark-centric bias in dataset selection (NQ, HotPotQA, TriviaQA, MS MARCO, Wikipedia); Selection bias from full-text availability requirement; Citation-lag bias from lower citation thresholds for 2025 publications; Time-window bias from 2020-2025 restriction; Language bias from English-only inclusion; Database coverage bias despite using five major sources; Selection bias toward highly-cited works; Domain skew bias: nearly half of studies concentrate on knowledge-intensive and open-domain QA settings; Citation-lag bias: Application of lower citation thresholds (≥15) to 2025 publications to mitigate inclusion bias, but acknowledged that this may still affect coverage; Time-window bias: Restriction to 2020–2025 may favor earlier, highly-cited papers; Language bias: English-only inclusion criterion may exclude relevant non-English literature; Database coverage bias: Five-library search strategy may miss studies in other venues or repositories; Domain skew bias: Evidence concentrated in knowledge-intensive and open-domain QA tasks, limiting transferability to other domains; Benchmark-centric bias: Literature appears optimized for benchmark achievement rather than real-world deployment metrics; Selection bias in data extraction: Single reviewer performed initial data extraction using RAG framework, though human verification was performed; RAG framework poses risks of hallucination and missing key data

Limitations

  • The authors acknowledge several key limitations: "We note the evidence base may be affected by citation-lag from the inclusion thresholds and by English-only, five-library coverage." Additionally, "Restricting the review to 2020–2025 could introduce a time-window bias by favouring earlier, highly cited papers." The review also notes that "coverage is still uneven, which limits how confidently methods can be lifted from benchmark QA and dropped into domain-specific pipelines without additional tuning," particularly in areas with scattered studies (finance, education, security, biomedical) where "evaluation practices are heterogeneous and often under-specified, making cross-paper comparisons fragile." The authors further state that "many reported gains emphasise retrieval–generation mechanics (e.g., retriever choice, reranking, context shaping) over deployment concerns," and "evidence is thinner for low-resource or safety-critical contexts where privacy, governance, or drift matter as much as accuracy."

Open questions raised

  • Inconsistent evaluation practices across domains limiting cross-paper comparisons
  • Thin evidence base for safety-critical, low-resource, and privacy-governed contexts
  • Lack of standardised reporting of system-level metrics (cost per query, end-to-end latency)
  • Uneven coverage in applied domains (finance, education, security, biomedical)
  • Need for joint reporting of accuracy, cost, robustness, and security in single studies
  • Missing data on parameter-efficient fine-tuning effects on RAG outcomes
Data: Structured database available via Zenodo (DOI: 10.5281/zenodo.17339384; uploaded 13 October 2025); Extracted data made publicly available via Zenodo (DOI: 10.5281/zenodo.17339384; uploaded 13 October 2025); Natural Questions (NQ); HotPotQA; TriviaQA; MS MARCO; Wikipedia; FinanceBench; Extracted study-level data publicly available via Zenodo (DOI: 10.5281/zenodo.17339384; uploaded 13 October 2025)Code: Python 3.12 scripts for deduplication and export (stated to be made available during peer review and prior to publication); All scripts used for deduplication and export will be made available during peer review and prior to publication; All Python scripts used for deduplication and export stated to be made available during peer review and prior to publication (specific repository URLs not provided in abstract/introduction)Extracted from: pdfAgreement 74%

Explore related topics

Related papers