A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges
Big Data and Cognitive Computing · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/bdcc9120320
Methodology & findings
Study design
Systematic literature review conducted according to PRISMA 2020 guidelines.
Primary method
Systematic literature review following PRISMA 2020
Main result
The review synthesized 128 studies on retrieval-augmented generation (RAG) and found that "methods have shifted from DPR+seq2seq baselines to modular, policy-driven RAG with hybrid/structure-aware retrieval, uncertainty-triggered loops, memory, and emerging multimodality." Additionally, "evaluation remains overlap-heavy (EM/F1), with increasing use of retrieval diagnostics (e.g., Recall@k, MRR@k), human judgements, and LLM-as-judge protocols." The evidence base shows that "nearly half of all studies concentrating on knowledge-intensive and open-domain QA settings," with emerging applications in software engineering (10.2%) and medical domains (8.6%).
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that "Evidence supports a shift to modular, policy-driven RAG, combining hybrid/structure-aware retrieval, uncertainty-aware control, memory, and multimodality, to improve grounding and efficiency." They further recommend: "(i) holistic benchmarks pairing quality with cost/latency and safety, (ii) budget-aware retrieval/tool-use policies, and (iii) provenance-aware pipelines that expose uncertainty and deliver traceable evidence." They note that "To advance from prototypes to dependable systems" these recommendations are essential for the field's maturation.
Risk of bias
Citation-lag bias from applying lower citation thresholds to 2025 publications; Time-window bias from 2020-2025 restriction favouring earlier, highly cited papers; Language bias (English-only inclusion); Database coverage bias (five libraries only); Domain concentration bias (nearly 50% of studies in QA/knowledge-intensive tasks); Benchmark-centric bias in dataset selection (NQ, HotPotQA, TriviaQA, MS MARCO, Wikipedia); Selection bias from full-text availability requirement; Citation-lag bias from lower citation thresholds for 2025 publications; Time-window bias from 2020-2025 restriction; Language bias from English-only inclusion; Database coverage bias despite using five major sources; Selection bias toward highly-cited works; Domain skew bias: nearly half of studies concentrate on knowledge-intensive and open-domain QA settings; Citation-lag bias: Application of lower citation thresholds (≥15) to 2025 publications to mitigate inclusion bias, but acknowledged that this may still affect coverage; Time-window bias: Restriction to 2020–2025 may favor earlier, highly-cited papers; Language bias: English-only inclusion criterion may exclude relevant non-English literature; Database coverage bias: Five-library search strategy may miss studies in other venues or repositories; Domain skew bias: Evidence concentrated in knowledge-intensive and open-domain QA tasks, limiting transferability to other domains; Benchmark-centric bias: Literature appears optimized for benchmark achievement rather than real-world deployment metrics; Selection bias in data extraction: Single reviewer performed initial data extraction using RAG framework, though human verification was performed; RAG framework poses risks of hallucination and missing key data
Limitations
- The authors acknowledge several key limitations: "We note the evidence base may be affected by citation-lag from the inclusion thresholds and by English-only, five-library coverage." Additionally, "Restricting the review to 2020–2025 could introduce a time-window bias by favouring earlier, highly cited papers." The review also notes that "coverage is still uneven, which limits how confidently methods can be lifted from benchmark QA and dropped into domain-specific pipelines without additional tuning," particularly in areas with scattered studies (finance, education, security, biomedical) where "evaluation practices are heterogeneous and often under-specified, making cross-paper comparisons fragile." The authors further state that "many reported gains emphasise retrieval–generation mechanics (e.g., retriever choice, reranking, context shaping) over deployment concerns," and "evidence is thinner for low-resource or safety-critical contexts where privacy, governance, or drift matter as much as accuracy."
Open questions raised
- Inconsistent evaluation practices across domains limiting cross-paper comparisons
- Thin evidence base for safety-critical, low-resource, and privacy-governed contexts
- Lack of standardised reporting of system-level metrics (cost per query, end-to-end latency)
- Uneven coverage in applied domains (finance, education, security, biomedical)
- Need for joint reporting of accuracy, cost, robustness, and security in single studies
- Missing data on parameter-efficient fine-tuning effects on RAG outcomes
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations
- Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studiesRuiqi Deng · 2024 · 295 citations