Automated systematic reviews using machine learning and large language models in clinical practice guideline development: A perspective
Takehiko Oami, Yohei Okada, Taka‐aki Nakada · Hong Kong Journal of Emergency Medicine · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/hkj2.70085
Methodology & findings
Study design
Perspective article describing prospective evaluation of semi-automated citation screening using machine learning (ASReview) and large language models (GPT-4 Turbo) tested on five clinical questions from Japanese Clinical Practice Guidelines for Management of Sepsis and Septic Shock 2024 (J-SSCG2024).
Sample
< 30
Primary method
The paper reports sensitivity, specificity, positive predictive value, and screening time comparisons. Confidence intervals are calculated but specific statistical tests are not explicitly detailed. The study appears to use descriptive statistics and comparisons across groups (different clinical questions and different LLM models). Reference standards were used to calculate test performance metrics.
Main result
The study found that "screening time was markedly reduced with the semi-automated method (1.3 min per 100 records) compared with the conventional method (17.2 min per 100 records)." For LLM-assisted screening, "LLM-assisted screening achieved a sensitivity of 0.75 (95% confidence interval [CI], 0.43-0.92) and specificity of 0.99 (95% CI, 0.99-0.99)," and "the LLM-assisted method reduced this time to 1.3 min (mean difference -15.25 min and 95% CI, -17.70 to -12.79)." With refined prompts, "sensitivity increased to 0.91 (95% CI, 0.77-0.97), whereas specificity remained high at 0.98 (95% CI, 0.96-0.99)."
Reports effect sizes and confidence intervals.
Research paradigm
Pragmatist/Applied research - evaluating feasibility and efficiency of automation tools in clinical practice guideline development
Author conclusions
The authors conclude that "automation tools have substantial potential to enhance efficiency and reduce the burden of the SR process. As AI capabilities continue to advance, further improvements in accuracy and reliability are expected through prompt engineering, domain adaptation, and model refinement. Thoughtful integration of these tools into guideline development workflows may support the sustainable production of high-quality clinical practice guidelines, thereby strengthening evidence-based decision-making and ultimately improving patient outcomes." They emphasize that "human experts must remain responsible for supervising AI outputs, validating evidence interpretations, and ensuring that recommendations remain clinically relevant and trustworthy."
Risk of bias
Selection bias in choice of test cases (5 CQs from sepsis domain may not generalize); Reference standard variation (gold standard defined differently across analyses); Domain-specific performance variability not controlled across clinical specialties; Prompt engineering dependency introduces subjective bias in model configuration; Potential publication bias in reported LLM performance studies (studies reporting negative results may be underrepresented); Variable performance across clinical domains - performance may not generalize beyond sepsis/critical care; Reviewer expertise differences affecting semi-automated tool performance; Domain-specific characteristics of literature affecting variability; Reference standard selection bias - sensitivity/specificity vary depending on whether gold standard is conventional screening vs. final inclusion list; Prompt design sensitivity - small changes in wording alter sensitivity/specificity; Model version dependency - newer versions may not show task-specific improvements; False negative rates from LLM classifications; PDF parsing failures affecting data extraction accuracy; Selection bias in choice of test cases: only five clinical questions related to initial resuscitation selected as test cases, potentially not representative of all guideline development scenarios; Domain specificity bias: evaluation focused on sepsis and critical care research, which may not generalize to other clinical specialties; Reference standard variability: different reference standards used in primary analysis (conventional title/abstract screening) versus secondary analysis (final inclusion list) affected performance metrics; Reviewer expertise variability: performance metrics may reflect differences in reviewer expertise and domain-specific characteristics; Model-specific performance variation: different LLM models showed substantially different performance characteristics, suggesting potential for selective reporting of favorable results; Prompt engineering bias: substantial improvements in sensitivity through prompt refinement suggest that initial performance assessments may underestimate model capabilities with optimization
Limitations
- The authors state that "the observed variability in performance highlights the limitations of relying on automated screening alone and suggests a need for further refinement and the importance of human oversight in guideline development." They further note that "model performance remains variable across clinical domains
- This issue requires careful validation whenever applying ML or LLMs to new topics." Additionally, "interpretability remains a major barrier
- The 'black-box' nature of advanced models can lead to hesitancy among clinicians and guideline committee members in the early phase of implementation." They also note that "many platforms require basic programming skills, knowledge of APIs, or local server integration
- These requirements may limit adoption by reviewers without technical expertise." Furthermore, "prompt design is critical to obtain accurate and consistent outputs
- Even small changes in wording may alter sensitivity or specificity."
Open questions raised
- Optimal strategy for integrating automated processes into guideline development not clearly established
- Effective implementation of automated SR tasks requires further research
- Appropriate balance between AI assistance and human judgment remains unclear
- Methodological and ethical considerations for safe and transparent adoption of automation need clarification
- Enhancements in accuracy and consistency across diverse clinical domains required
- Full automation remains infeasible for complex tasks like ROB assessment (median error rate of 27%)
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations