12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Hybrid Delphi-Inspired Expert–LLM Workflow for Efficient Evidence Screening in Systematic Reviews

Omid Pournik, Emma Watts, Emma Richards, Kristien Boelaert, Neil Sharma, Saadullah Farooq Abbasi et al. · Studies in health technology and informatics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3233/shti260255

Methodology & findings

Study design

Observational study with workflow evaluation.

Sample

N = 14858, 6 groups

Primary method

Cohen's kappa (κ) for inter-rater agreement; concordance percentage calculation. No other statistical methods or software are specified in the abstract.

Main result

The workflow achieved "96% concordance (κ = 0.91) with human reviewers, with only one false exclusion, and reduced manual screening time by approximately 70%", demonstrating that the hybrid expert-LLM approach can accurately automate early-stage evidence screening while maintaining high performance standards.

Reports effect sizes.

Research paradigm

Positivist/empiricist with pragmatic AI integration

Author conclusions

The authors conclude that "a transparent Delphi-inspired expert-LLM can accurately and reproducibly automate early-stage evidence screening, providing substantial efficiency gains while preserving human oversight and methodological rigor" and that "the approach offers a practical pathway toward the responsible integration of generative AI in systematic review methodology and digital health research."

Risk of bias

Selection bias in random sample validation (only 100 of 14,858 records independently reviewed); Potential model-specific bias (ChatGPT-5 training data and design); Limited generalizability (domain-specific to thyroid nodule malignancy risk assessment); Limited sample validation (only 100 of 14,858 records independently reviewed by humans - 0.67%); Single model tested (ChatGPT-5 only - no comparison with other LLMs); Potential selection bias in which 100 records were randomly sampled; No information on blinding of human reviewers to LLM classifications; Limited demographic information on expert panel composition; Study domain-specific (thyroid nodule risk) - generalizability unclear; Selection bias: Only 100 records (0.7%) randomly sampled for human validation; Potential funding bias: ChatGPT-5 used as primary tool; no disclosure of OpenAI funding or conflicts; Single false exclusion not contextualized relative to false inclusions; No assessment of systematic errors in specific document types or domains

Open questions raised

  • The authors identify the need for broader validation across different systematic review contexts and the importance of maintaining human oversight in AI-integrated evidence screening workflows.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 61%

Explore related topics

Related papers