Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in Psychiatry: Nouvel algorithme de sélection des résumés utilisant un consensus issu d’un grand modèle de langage inspiré de la méthode Delphi pour les revues systématiques en psychiatrie
Mirkamal Tolend, Ramzi Halabi, Kousai Ghaouari, Yvonne C Y Lau, Martin Alda, Arend Hintze et al. · The Canadian Journal of Psychiatry · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1177/07067437261445767
Methodology & findings
Study design
Algorithm development and validation study using three published systematic reviews in psychiatry.
Sample
N = 12039, 3 groups
Primary method
Logistic regression model for ranking abstracts; Delphi-inspired consensus process with ensemble of five LLMs; probability threshold-based filtering; recall and work saved at 95% recall (WSS@95%) as evaluation metrics.
Main result
The study found that the Delphi-LLM workflow "correctly identified 1,605 (97.0%) of these 1,655 abstracts (precision = 54.2%, WSS@95% = 38.1%)" on the autism biomarkers dataset, and "The recall and work saved metrics were similarly reliable and among the top in two low-prevalence datasets on an attention-deficit hyperactivity disorder treatment review (10% of 2,891 relevant) and a posttraumatic stress disorder trajectory review (7% of 4,453 relevant). For these two datasets, recall was 100.0% and 96.4%, and the WSS@95% was 17.3% and 18.5%, respectively."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that they "presented the design and validation of a novel abstract screening workflow that centres around a Delphi-style aggregation process to harness the strengths of five open-source LLMs that can be run on consumer-level workstations. This multi-LLM workflow showed acceptable and reliable performance for use as an automated prescreening method to facilitate systematic reviews."
Risk of bias
Selection bias: validation performed only on previously screened abstracts from published reviews; Potential dependency on training data from seed examples; LLM hallucination or inconsistency risks not discussed; Limited to psychiatry domain and three specific review topics; Selection bias: The workflow was tested on abstracts already screened in published systematic reviews, potentially biasing results based on the original authors' screening decisions; Lack of blinding: Authors who created eligibility criteria and seed examples may have influenced LLM classification; Limited scope: Testing limited to three psychiatric systematic reviews; generalizability to other domains unclear
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations