The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review
Dmitry Scherbakov, Nina Hubig, Vinita Jansari, Alexander Bakumenko, Leslie Lenert · Journal of the American Medical Informatics Association · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1093/jamia/ocaf063
Methodology & findings
Study design
Systematic review with LLM-assisted automation.
Sample
N = 172, 9 groups
Primary method
Frequency count as main synthesis method; boxplots for displaying median, interquartile range (IQR), whiskers (within 1.5 × IQR), and outliers; mean and standard deviation for comparing performance metrics; column plots for presenting frequencies; geographic mapping for location data. Studies reporting only qualitative metrics or metrics other than those specified in the extraction form were not included in quantitative comparisons.
Main result
The study found that "GPT models had lower accuracy in title/abstract screening (M = 77.34, SD = 13.06) compared to BERT models (M = 80.87, SD = 11.81). However, GPT models performed better in data extraction, with precision (M = 83.07, SD = 10.43) and recall (M = 85.99, SD = 9.82), while BERT models had lower precision (M = 61.06, SD = 31.26)" Additionally, "The use of LLMs in review automation is rapidly growing, with expected radical changes in scientific evidence synthesis. LLMs are likely to significantly reduce the time needed for reviews while producing similar or higher-quality data in greater quantities than manual reviews do."
Reports effect sizes.
Research paradigm
Mixed methods (quantitative systematic review with qualitative assessment)
Author conclusions
"The use of LLMs in review automation is rapidly growing, with expected radical changes in scientific evidence synthesis. LLMs are likely to significantly reduce the time needed for reviews while producing similar or higher-quality data in greater quantities than manual reviews do." However, "Despite early successes, few systematic reviews using LLMs were identified in our review. Although still in its early stages, AI-assisted reviews are already yielding impressive results, with growing interest as researchers develop semi-automated pipelines. However, generating trustworthy and useful AI-driven reviews still presents both technological and ethical challenges, particularly for quantitative meta-analyses comparing treatment effects."
Risk of bias
Publication bias: reliance on disclosed LLM usage only; potential missed studies using LLMs but not disclosing; Selection bias: English-language only inclusion; Extraction accuracy bias: performance metrics extraction had lower accuracy (<0.8 precision); Automation bias: use of LLM as reviewer may introduce systematic biases; Funding bias: majority (56.4%) had public funding, potentially underrepresenting industry-funded studies; Dependence on author disclosure of LLM usage without automated detection; No automatic LLM usage detection method employed; Potential selection bias toward published studies; Lower accuracy in performance metrics extraction category; Variable quality of extracted data across heterogeneous study designs; Potential publication bias toward positive results; Selection bias: Reliance on author disclosure of LLM usage; undisclosed LLM-generated reviews may have been missed; Extraction bias: LLM-based extraction with lower accuracy in some categories (acknowledged as <80% precision in some fields); Publication bias: Only English-language journal publications included; conference abstracts and non-English publications excluded; Automation bias: As noted in introduction, reliance on automated systems could lead to overlooked errors; No independent quality assessment performed due to diverse publication types
Limitations
- "Some extraction categories, such as performance metrics, had relatively lower accuracy
- Therefore, the results of this extraction category should be taken with caution." Additionally, "we relied on the disclosure of LLM usage by the authors of reviewed publications, and this study did not use any type of automatic LLM usage detection
- thus, we could have missed publications, especially potential reviews, that could have been created with LLM support." The authors also note that "Due to the diverse nature of publications and study designs, bias and quality assessment were not performed."
Open questions raised
- Few studies focus on full cycle review automation; most address specific areas like extraction or screening
- Limited research on integrating and evaluating different LLMs in combination
- Long-term impacts and ethical considerations of LLM integration need assessment
- Improved LLM performance metrics, particularly precision and recall in lower-accuracy extraction categories
- Research on impact of LLM-generated content on journal review processes and publication barriers
- Development of LLM-assisted approaches for quantitative meta-analyses
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations