12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automated systematic reviews using machine learning and large language models in clinical practice guideline development: A perspective

Takehiko Oami, Yohei Okada, Taka‐aki Nakada · Hong Kong Journal of Emergency Medicine · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
2
Citations
28.34
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/hkj2.70085

Methodology & findings

Study design

Perspective article describing prospective evaluation of semi-automated citation screening using machine learning (ASReview) and large language models (GPT-4 Turbo) tested on five clinical questions from Japanese Clinical Practice Guidelines for Management of Sepsis and Septic Shock 2024 (J-SSCG2024).

Sample

< 30

Primary method

The paper reports sensitivity, specificity, positive predictive value, and screening time comparisons. Confidence intervals are calculated but specific statistical tests are not explicitly detailed. The study appears to use descriptive statistics and comparisons across groups (different clinical questions and different LLM models). Reference standards were used to calculate test performance metrics.

Main result

The study found that "screening time was markedly reduced with the semi-automated method (1.3 min per 100 records) compared with the conventional method (17.2 min per 100 records)." For LLM-assisted screening, "LLM-assisted screening achieved a sensitivity of 0.75 (95% confidence interval [CI], 0.43-0.92) and specificity of 0.99 (95% CI, 0.99-0.99)," and "the LLM-assisted method reduced this time to 1.3 min (mean difference -15.25 min and 95% CI, -17.70 to -12.79)." With refined prompts, "sensitivity increased to 0.91 (95% CI, 0.77-0.97), whereas specificity remained high at 0.98 (95% CI, 0.96-0.99)."

Reports effect sizes and confidence intervals.

Research paradigm

Pragmatist/Applied research - evaluating feasibility and efficiency of automation tools in clinical practice guideline development

Author conclusions

The authors conclude that "automation tools have substantial potential to enhance efficiency and reduce the burden of the SR process. As AI capabilities continue to advance, further improvements in accuracy and reliability are expected through prompt engineering, domain adaptation, and model refinement. Thoughtful integration of these tools into guideline development workflows may support the sustainable production of high-quality clinical practice guidelines, thereby strengthening evidence-based decision-making and ultimately improving patient outcomes." They emphasize that "human experts must remain responsible for supervising AI outputs, validating evidence interpretations, and ensuring that recommendations remain clinically relevant and trustworthy."

Risk of bias

Selection bias in choice of test cases (5 CQs from sepsis domain may not generalize); Reference standard variation (gold standard defined differently across analyses); Domain-specific performance variability not controlled across clinical specialties; Prompt engineering dependency introduces subjective bias in model configuration; Potential publication bias in reported LLM performance studies (studies reporting negative results may be underrepresented); Variable performance across clinical domains - performance may not generalize beyond sepsis/critical care; Reviewer expertise differences affecting semi-automated tool performance; Domain-specific characteristics of literature affecting variability; Reference standard selection bias - sensitivity/specificity vary depending on whether gold standard is conventional screening vs. final inclusion list; Prompt design sensitivity - small changes in wording alter sensitivity/specificity; Model version dependency - newer versions may not show task-specific improvements; False negative rates from LLM classifications; PDF parsing failures affecting data extraction accuracy; Selection bias in choice of test cases: only five clinical questions related to initial resuscitation selected as test cases, potentially not representative of all guideline development scenarios; Domain specificity bias: evaluation focused on sepsis and critical care research, which may not generalize to other clinical specialties; Reference standard variability: different reference standards used in primary analysis (conventional title/abstract screening) versus secondary analysis (final inclusion list) affected performance metrics; Reviewer expertise variability: performance metrics may reflect differences in reviewer expertise and domain-specific characteristics; Model-specific performance variation: different LLM models showed substantially different performance characteristics, suggesting potential for selective reporting of favorable results; Prompt engineering bias: substantial improvements in sensitivity through prompt refinement suggest that initial performance assessments may underestimate model capabilities with optimization

Limitations

  • The authors state that "the observed variability in performance highlights the limitations of relying on automated screening alone and suggests a need for further refinement and the importance of human oversight in guideline development." They further note that "model performance remains variable across clinical domains
  • This issue requires careful validation whenever applying ML or LLMs to new topics." Additionally, "interpretability remains a major barrier
  • The 'black-box' nature of advanced models can lead to hesitancy among clinicians and guideline committee members in the early phase of implementation." They also note that "many platforms require basic programming skills, knowledge of APIs, or local server integration
  • These requirements may limit adoption by reviewers without technical expertise." Furthermore, "prompt design is critical to obtain accurate and consistent outputs
  • Even small changes in wording may alter sensitivity or specificity."

Open questions raised

  • Optimal strategy for integrating automated processes into guideline development not clearly established
  • Effective implementation of automated SR tasks requires further research
  • Appropriate balance between AI assistance and human judgment remains unclear
  • Methodological and ethical considerations for safe and transparent adoption of automation need clarification
  • Enhancements in accuracy and consistency across diverse clinical domains required
  • Full automation remains infeasible for complex tasks like ROB assessment (median error rate of 27%)
Extracted from: pdfAgreement 59%

Explore related topics

Related papers