12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The emergence of large language models as tools in literature reviews: a large language model-assisted systematic review

Dmitry Scherbakov, Nina Hubig, Vinita Jansari, Alexander Bakumenko, Leslie Lenert · Journal of the American Medical Informatics Association · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
0/4
Quality (LMQS)
E
Evidence
95
Citations
86.23
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1093/jamia/ocaf063

Methodology & findings

Study design

Systematic review with LLM-assisted automation.

Sample

N = 172, 9 groups

Primary method

Frequency count as main synthesis method; boxplots for displaying median, interquartile range (IQR), whiskers (within 1.5 × IQR), and outliers; mean and standard deviation for comparing performance metrics; column plots for presenting frequencies; geographic mapping for location data. Studies reporting only qualitative metrics or metrics other than those specified in the extraction form were not included in quantitative comparisons.

Main result

The study found that "GPT models had lower accuracy in title/abstract screening (M = 77.34, SD = 13.06) compared to BERT models (M = 80.87, SD = 11.81). However, GPT models performed better in data extraction, with precision (M = 83.07, SD = 10.43) and recall (M = 85.99, SD = 9.82), while BERT models had lower precision (M = 61.06, SD = 31.26)" Additionally, "The use of LLMs in review automation is rapidly growing, with expected radical changes in scientific evidence synthesis. LLMs are likely to significantly reduce the time needed for reviews while producing similar or higher-quality data in greater quantities than manual reviews do."

Reports effect sizes.

Research paradigm

Mixed methods (quantitative systematic review with qualitative assessment)

Author conclusions

"The use of LLMs in review automation is rapidly growing, with expected radical changes in scientific evidence synthesis. LLMs are likely to significantly reduce the time needed for reviews while producing similar or higher-quality data in greater quantities than manual reviews do." However, "Despite early successes, few systematic reviews using LLMs were identified in our review. Although still in its early stages, AI-assisted reviews are already yielding impressive results, with growing interest as researchers develop semi-automated pipelines. However, generating trustworthy and useful AI-driven reviews still presents both technological and ethical challenges, particularly for quantitative meta-analyses comparing treatment effects."

Risk of bias

Publication bias: reliance on disclosed LLM usage only; potential missed studies using LLMs but not disclosing; Selection bias: English-language only inclusion; Extraction accuracy bias: performance metrics extraction had lower accuracy (<0.8 precision); Automation bias: use of LLM as reviewer may introduce systematic biases; Funding bias: majority (56.4%) had public funding, potentially underrepresenting industry-funded studies; Dependence on author disclosure of LLM usage without automated detection; No automatic LLM usage detection method employed; Potential selection bias toward published studies; Lower accuracy in performance metrics extraction category; Variable quality of extracted data across heterogeneous study designs; Potential publication bias toward positive results; Selection bias: Reliance on author disclosure of LLM usage; undisclosed LLM-generated reviews may have been missed; Extraction bias: LLM-based extraction with lower accuracy in some categories (acknowledged as <80% precision in some fields); Publication bias: Only English-language journal publications included; conference abstracts and non-English publications excluded; Automation bias: As noted in introduction, reliance on automated systems could lead to overlooked errors; No independent quality assessment performed due to diverse publication types

Limitations

  • "Some extraction categories, such as performance metrics, had relatively lower accuracy
  • Therefore, the results of this extraction category should be taken with caution." Additionally, "we relied on the disclosure of LLM usage by the authors of reviewed publications, and this study did not use any type of automatic LLM usage detection
  • thus, we could have missed publications, especially potential reviews, that could have been created with LLM support." The authors also note that "Due to the diverse nature of publications and study designs, bias and quality assessment were not performed."

Open questions raised

  • Few studies focus on full cycle review automation; most address specific areas like extraction or screening
  • Limited research on integrating and evaluating different LLMs in combination
  • Long-term impacts and ethical considerations of LLM integration need assessment
  • Improved LLM performance metrics, particularly precision and recall in lower-accuracy extraction categories
  • Research on impact of LLM-generated content on journal review processes and publication barriers
  • Development of LLM-assisted approaches for quantitative meta-analyses
Extracted from: pdfAgreement 54%

Explore related topics

Related papers