Large language models for conducting systematic reviews: on the rise, but not yet ready for use—a scoping review
Judith-Lisa Lieberum, Maria‐Inti Metzendorf, Felix Heilmeyer, Waldemar Siemens, Christian Haverkamp, Daniel Böhringer et al. · Journal of Clinical Epidemiology · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.jclinepi.2025.111746
Methodology & findings
Study design
Scoping review with systematic search across multiple databases (MEDLINE, Web of Science, IEEEXplore, ACM Digital Library, Europe PMC, Google Scholar) and hand search.
Main result
The study found that "LLM approaches covered 10 of 13 defined SR steps, most frequently literature search (n = 15, 41%), study selection (n = 14, 38%), and data extraction (n = 11, 30%)" with "the mostly recurring LLM was Generative Pretrained Transformer (GPT) (n = 33, 89%)". Furthermore, "in half of the studies, authors evaluated LLM use as promising (n = 20, 54%), one-quarter as neutral (n = 9, 24%) and one-fifth as nonpromising (n = 8, 22%)".
Research paradigm
Systematic evidence synthesis / empirical analysis of published literature
Author conclusions
The authors conclude that "Although LLMs show promise in supporting SR creation, fully established or validated applications are often lacking" and that "The rapid increase in research on LLMs for evidence synthesis production highlights their growing relevance."
Risk of bias
Language restriction (English or German only); Publication bias (searches limited to published articles and preprints); Time-period bias (studies from April 2021 onwards only); Database selection bias (searches across specific databases may miss relevant grey literature)
Limitations
- Although not explicitly stated as limitations in the abstract, the authors note that "fully established or validated applications are often lacking," suggesting a key limitation is the premature state of LLM applications for systematic review conduct
- The predominance of "validation studies (n = 21, 57%)" over implementation studies indicates a limitation in the maturity of the evidence base.
Open questions raised
- The authors identify that LLMs lack full validation and established applications for systematic review conduct. The growing number of studies suggests increasing relevance but indicates the field is still developing.
- The authors identify that fully established or validated applications of LLMs for systematic reviews are often lacking, suggesting a need for more rigorous validation and development of LLM tools for SR conduct.
- The authors identify that while LLMs show potential for SR support, there is "a lack of fully tested and validated applications," suggesting need for more rigorous validation and implementation studies before clinical or research adoption.
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations