A Hybrid Delphi-Inspired Expert–LLM Workflow for Efficient Evidence Screening in Systematic Reviews
Omid Pournik, Emma Watts, Emma Richards, Kristien Boelaert, Neil Sharma, Saadullah Farooq Abbasi et al. · Studies in health technology and informatics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3233/shti260255
Methodology & findings
Study design
Observational study with workflow evaluation.
Sample
N = 14858, 6 groups
Primary method
Cohen's kappa (κ) for inter-rater agreement; concordance percentage calculation. No other statistical methods or software are specified in the abstract.
Main result
The workflow achieved "96% concordance (κ = 0.91) with human reviewers, with only one false exclusion, and reduced manual screening time by approximately 70%", demonstrating that the hybrid expert-LLM approach can accurately automate early-stage evidence screening while maintaining high performance standards.
Reports effect sizes.
Research paradigm
Positivist/empiricist with pragmatic AI integration
Author conclusions
The authors conclude that "a transparent Delphi-inspired expert-LLM can accurately and reproducibly automate early-stage evidence screening, providing substantial efficiency gains while preserving human oversight and methodological rigor" and that "the approach offers a practical pathway toward the responsible integration of generative AI in systematic review methodology and digital health research."
Risk of bias
Selection bias in random sample validation (only 100 of 14,858 records independently reviewed); Potential model-specific bias (ChatGPT-5 training data and design); Limited generalizability (domain-specific to thyroid nodule malignancy risk assessment); Limited sample validation (only 100 of 14,858 records independently reviewed by humans - 0.67%); Single model tested (ChatGPT-5 only - no comparison with other LLMs); Potential selection bias in which 100 records were randomly sampled; No information on blinding of human reviewers to LLM classifications; Limited demographic information on expert panel composition; Study domain-specific (thyroid nodule risk) - generalizability unclear; Selection bias: Only 100 records (0.7%) randomly sampled for human validation; Potential funding bias: ChatGPT-5 used as primary tool; no disclosure of OpenAI funding or conflicts; Single false exclusion not contextualized relative to false inclusions; No assessment of systematic errors in specific document types or domains
Open questions raised
- The authors identify the need for broader validation across different systematic review contexts and the importance of maintaining human oversight in AI-integrated evidence screening workflows.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations