Automated approaches to identifying clinical trials based on title and abstract in the field of physiotherapy: a comparative analysis
Wing S Kwok, Geraldine Wallbank, Philip Hodgson, Thomas Schräder, Lexuan Shao, Mark Elkins et al. · Journal of Clinical Epidemiology · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.jclinepi.2026.112309
Methodology & findings
Study design
Comparative performance evaluation study.
Sample
N = 10793, 7 groups
Primary method
Classification metrics evaluated: accuracy, precision, recall, and F1 score (harmonic mean of precision and recall), all reported as percentages with 95% confidence intervals. Agreement between raters assessed using Cohen's kappa. Time efficiency compared between human and LLM approaches. Data analyzed using STATA 16.0 (StataCorp LLC). Exploratory analyses used support vector machine (SVM), logistic regression, BERT-based natural language processing, and API-based retrieval-augmented generation (RAG) implementation.
Main result
The study found that "the accuracy of ChatGPT and Copilot in the initial test ranged from 83% to 87%" with Copilot achieving higher precision (51%) but lower recall (70%) than ChatGPT. The research demonstrated that "commercial web-based LLMs have the potential to support title and abstract screening to identify relevant records," though "alternative approaches, such as machine learning and natural language processing, could achieve higher agreement." Importantly, "the LLM-based approach required 37% of the time of the human-only screening process."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist (computational performance evaluation against reference standard)
Author conclusions
"This study demonstrates the potential of commercial, web-based LLM tools to support title and abstract screening and improve efficiency of evidence syntheses. Alternative approaches to automation, such as machine learning and natural language processing, could achieve screening performance similar to or slightly higher than that of commercial tools, but they require a series of preprocessing steps." The authors further concluded that "Commercial, web-based LLMs have the potential to support title and abstract screening to identify relevant records" but that "they may be useful as a filter, eliminating most irrelevant records and substantially reducing the initial human screening burden, although manual review remains necessary."
Risk of bias
Reference standard bias: Human reviewers may have decision-making errors despite being the 'reference standard'; In-context learning bias: Use of previously screened articles for training may introduce misleading patterns; Model instability: Commercial LLM tools showed temporal variation in performance between initial and repeat testing; Prompt sensitivity: Results may be dependent on specific prompt wording used; Generalizability bias: Findings may not apply to other disciplines or broader physiotherapy field applications; Opaqueness of commercial LLM tools may introduce performance variability; Reference standard (human consensus) may involve decision-making error; In-context learning approach may introduce misleading information; Prompt wording may influence LLM output and limit generalizability; Potential temporal variability in commercial tools (model updates, deployment changes); Reference standard bias: human decision-making error in the reference standard could affect performance metrics; Selection bias: in-context learning may introduce misleading information influencing performance metrics; Model opacity: commercial web-based LLMs are 'black-box' systems with opaque architectures; Temporal variability: significant differences in performance between initial and repeat testing (especially Copilot); Prompt sensitivity: results depend on specific prompt wording which may not generalize; Interrater reliability concern: initial human interrater agreement was moderate (Cohen's kappa 0.48)
Limitations
- The authors acknowledged several limitations: "First, commercial, web-based LLM tools are opaque at the architectural level, and their performance may not be entirely stable." Additionally, "evaluation metrics depend on the reference standard and were used as in-context learning
- The reference standard may involve a degree of human decision-making error, and the use of in-context learning may inadvertently mean that misleading information may be provided." Furthermore, "the prompts used are tailored to this proof-of-concept study, and prompt wording may influence LLM output, and thus, the findings may not generalize other disciplines."
Open questions raised
- Lack of evaluation of LLMs for identifying relevant trials across the broader physiotherapy field beyond musculoskeletal specialty
- Limited evidence on temporal reproducibility and stability of commercial LLM performance
- Need for research on using LLMs to perform full-text article screening with richer contextual information
- Investigation of hybrid screening workflows combining LLM-based and traditional machine learning or NLP techniques as parallel reviewers
- Use of LLMs to identify duplicates among search results
- Application of these approaches to other disciplines and evidence synthesis databases
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations