12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Applications of Large Language Models in Medical Research: From Systematic Reviews to Clinical Studies

Eun Jeong Gong, Chang Seok Bang, Yong Seok Shin · Bioengineering · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/bioengineering13030365

Methodology & findings

Study design

Narrative review synthesis.

Primary method

The paper reviews multiple statistical approaches used across studies, including: sensitivity and specificity calculations, Cohen's kappa (κ) for agreement assessment (ranging from κ = 0.16 to κ = 0.73 for risk-of-bias assessments), accuracy metrics (e.g., 80-94% for data extraction), F1 scores for search strategy evaluation (ranging from 0.077 to 0.517), and various performance metrics across different LLM applications. No original statistical analysis was conducted by the authors.

Main result

The study found that "literature screening shows high sensitivity with substantial workload reduction, while tasks requiring subjective judgment, such as risk-of-bias assessment, remain insufficiently validated for standalone use." Additionally, "LLM-generated strategies frequently fail to incorporate synonymous entry terms, miss clinical practice jargon, incorrectly group acronyms, insert unjustified date limitations, and, critically, omit validated study design filters for identifying randomized controlled trials (RCTs)." For data extraction, "a collaborative dual-LLM approach using GPT-4-turbo and Claude-3-Opus significantly improves accuracy to 94% when both models agree, while reducing hallucination rates from approximately 2.5% to 0.25%."

Reports confidence intervals.

Research paradigm

interpretivist; mixed (systematic evidence synthesis with narrative synthesis of heterogeneous study designs)

Author conclusions

"LLMs are powerful but inherently unstable instruments requiring constant calibration-success depends on researchers maintaining their roles as critical overseers rather than passive consumers of AI-generated content. In practice, this means adopting iterative, step-by-step refinement rather than expecting polished output from single prompts, and rigorously verifying every AI-generated citation and claim against primary sources." The authors recommend a structured approach that "start[s] with low-risk applications, implement[s] multi-layered validation, maintain[s] reproducible settings, and preserve[s] human judgment for tasks requiring causal reasoning."

Risk of bias

Publication bias toward positive results may overestimate LLM capabilities; Heterogeneity of evaluation metrics across studies limits direct comparisons; Lack of transparency in proprietary LLM architectures and training data; Language bias: GPT-4 data extraction sensitivity drops from 75% for English articles to 36% for non-English publications; Demographic biases in medical LLMs: over 90% of studies identified demographic biases; Paywalled journal content not included in LLM training (subscription content from Elsevier, Springer Nature, Wiley); Clinical bias: GPT-4 more likely to rate Black patients as abusing opioids with identical clinical information; Publication bias favoring positive results may overestimate LLM capabilities; Selection bias from non-exhaustive literature search in narrative review format; Language bias: LLMs demonstrate significantly reduced performance on non-English texts (sensitivity drops from 75% for English to 36% for non-English publications); Demographic bias in medical LLMs: over 90% of studies identified demographic biases, including propagation of debunked race-based medicine; Reporting heterogeneity: inconsistent evaluation metrics across studies limit direct comparisons; Proprietary model bias: most evidence derives from non-transparent commercial models (GPT series); Publication bias toward positive results (authors explicitly note this may overestimate LLM capabilities); Language bias: documented performance degradation for non-English texts (sensitivity drops from 75% to 36% for non-English articles); Demographic bias: over 90% of studies identified demographic biases in medical LLMs; Selection bias in review: focus on studies with empirical performance data and validation metrics, potentially excluding null/negative findings; Model-specific bias: evidence predominantly from proprietary GPT-series models, with limited open-source model evaluation until recent developments

Limitations

  • This review has several limitations
  • As a narrative review rather than a systematic review, "our literature search and study selection, though structured, were not exhaustive
  • The rapid pace of LLM development means that some findings reviewed here may already be outdated
  • Publication bias toward positive results may overestimate LLM capabilities, and the heterogeneity of evaluation metrics across studies limits direct comparisons
  • Furthermore, most evidence derives from studies using proprietary commercial models (e.g., GPT-4), whose underlying architectures and training data are not fully transparent, limiting reproducibility and generalizability of findings."

Open questions raised

  • Lack of standardized benchmarks for evaluating LLM-generated search strategies against expert-crafted ones
  • Limited evaluation of model parameters (temperature settings) impact on screening accuracy
  • Insufficient validation of LLM-generated statistical analysis plans (SAPs) compared to human biostatisticians
  • Limited empirical validation of LLM applications in harmonizing writing styles across multi-author projects
  • Underexploration of net efficiency gains of LLM-assisted writing when accounting for verification burden
  • Limited rigorous validation comparing LLM-generated research questions with expert-derived hypotheses
Extracted from: pdfAgreement 60%

Explore related topics

Related papers