12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automating Systematic Literature Reviews with Retrieval-Augmented Generation: A Comprehensive Overview

Binglan Han, Teo Sušnjak, Anuradha Mathrani · Applied Sciences · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
3/4
Quality (LMQS)
I
Evidence
45
Citations
14.19
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/app14199103

Methodology & findings

Study design

Comprehensive narrative literature review examining Retrieval-Augmented Generation (RAG) in large language models and their applications to systematic literature reviews.

Main result

The study found that "RAG-based LLMs can facilitate and potentially automate these tasks to significantly improve efficiency and accuracy" in systematic literature reviews. The paper also identifies that "RAG mitigates hallucinations by anchoring outputs in verifiable sources, enabling users to trace and validate the information provided" and that "RAG-based LLMs have demonstrated considerable effectiveness in performing key tasks associated with SLRs, such as extracting critical data from scientific documents, generating concise summaries of individual papers and collections, providing accurate citation recommendations, and identifying emerging research trends."

Reports effect sizes.

Research paradigm

Interpretivist/qualitative review of technical literature

Author conclusions

The authors conclude: "The Retrieval-Augmented Generation (RAG) framework comprises three fundamental processes: retrieval, augmentation, and generation... RAG-based LLMs have significantly improved accuracy, relevance, and contextual comprehension by combining real-time information retrieval with generative functions. These enhancements position RAG-based LLMs as promising candidates for automating various tasks associated with systematic literature reviews." They further state: "Future research should aim to optimize the synergy between LLM selection, training strategies, RAG techniques, and prompt engineering to implement the proposed framework, with a particular focus on the retrieval of information from individual scientific papers and the integration of retrieved data to generate outputs on diverse aspects such as current status, existing gaps, and emerging trends."

Risk of bias

Selection bias: Limited empirical evidence from only three studies specifically focused on RAG-based LLMs for SLRs; Publication bias: Review only includes published/archived studies; gray literature not systematically searched; Recency bias: Heavy reliance on recent (2023-2024) research; limited historical perspective; Technology bias: Authors acknowledge using GPT-4 to strengthen sentence structure, creating potential circular reasoning regarding LLM capabilities; Scope limitation: Review does not comprehensively cover all RAG applications; selective focus on SLR-relevant tasks

Limitations

  • The authors note that "although all three studies utilize Retrieval-Augmented Generation (RAG)-based Large Language Models (LLMs) to enhance the Systematic Literature Review (SLR) process, they exhibit distinct methodologies and focal areas" and that "these studies are conducted on relatively small datasets or within limited scopes
  • Furthermore, the outputs generated by the frameworks discussed in these studies are typically brief responses to specific questions or they give relatively generalized summaries, which do not equate to a thorough critique that is conducted on selected articles." The authors also state: "there is a notable gap in datasets tailored for specific SLR tasks beyond general scientific article summarization."

Open questions raised

  • Integration of domain-specific LLMs in RAG systems for SLRs
  • Multimodal context for augmentation and multimodal output generation
  • Development of specialized datasets for specific SLR tasks beyond general scientific article summarization
  • Strategies for designing complete frameworks integrating RAG-based LLMs across all stages of SLR
  • Methods to improve data retrieval at varying levels of granularity from individual scientific papers
  • Techniques for more effective synthesis and integration of retrieved data to produce diverse outputs (summaries, meta-analyses, trend analyses)
Data: Multi-XScience - multi-document summarization dataset; ACLSum - aspect-based summarization dataset; MASSW - scientific workflows dataset; CHIME - hierarchical scientific studies dataset; Multi-XScience (large-scale dataset for multi-document summarization of scientific articles); ACLSum (expert-curated dataset for multi-aspect summarization of scientific papers); MASSW (dataset for summarizing multiple aspects of scientific workflows); CHIME (hierarchical dataset organizing scientific studies into tree structures); PubMed (academic papers repository); MEDLINE (medical research database); Wikipedia (general-purpose knowledge repository); Common Crawl (web crawl data repository); Wikidata (knowledge graph database); Semantic Scholar API (referenced for paper retrieval in LitLLM); PubMed API (referenced for paper metadata retrieval in RefAI); Multi-XScience - large-scale dataset for multi-document summarization of scientific articles; ACLSum - expert-curated dataset for multi-aspect summarization of scientific papers; MASSW - dataset for summarizing multiple aspects of scientific workflows; CHIME - hierarchical dataset organizing scientific studies into tree structures; PubMed - unstructured text corpus of academic papers; MEDLINE - domain-specific repository for medical research; Wikipedia - pre-trained knowledge base; Common Crawl - general-purpose datasetExtracted from: pdfAgreement 72%

Explore related topics

Related papers