12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automated approaches to identifying clinical trials based on title and abstract in the field of physiotherapy: a comparative analysis

Wing S Kwok, Geraldine Wallbank, Philip Hodgson, Thomas Schräder, Lexuan Shao, Mark Elkins et al. · Journal of Clinical Epidemiology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.jclinepi.2026.112309

Methodology & findings

Study design

Comparative performance evaluation study.

Sample

N = 10793, 7 groups

Primary method

Classification metrics evaluated: accuracy, precision, recall, and F1 score (harmonic mean of precision and recall), all reported as percentages with 95% confidence intervals. Agreement between raters assessed using Cohen's kappa. Time efficiency compared between human and LLM approaches. Data analyzed using STATA 16.0 (StataCorp LLC). Exploratory analyses used support vector machine (SVM), logistic regression, BERT-based natural language processing, and API-based retrieval-augmented generation (RAG) implementation.

Main result

The study found that "the accuracy of ChatGPT and Copilot in the initial test ranged from 83% to 87%" with Copilot achieving higher precision (51%) but lower recall (70%) than ChatGPT. The research demonstrated that "commercial web-based LLMs have the potential to support title and abstract screening to identify relevant records," though "alternative approaches, such as machine learning and natural language processing, could achieve higher agreement." Importantly, "the LLM-based approach required 37% of the time of the human-only screening process."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist (computational performance evaluation against reference standard)

Author conclusions

"This study demonstrates the potential of commercial, web-based LLM tools to support title and abstract screening and improve efficiency of evidence syntheses. Alternative approaches to automation, such as machine learning and natural language processing, could achieve screening performance similar to or slightly higher than that of commercial tools, but they require a series of preprocessing steps." The authors further concluded that "Commercial, web-based LLMs have the potential to support title and abstract screening to identify relevant records" but that "they may be useful as a filter, eliminating most irrelevant records and substantially reducing the initial human screening burden, although manual review remains necessary."

Risk of bias

Reference standard bias: Human reviewers may have decision-making errors despite being the 'reference standard'; In-context learning bias: Use of previously screened articles for training may introduce misleading patterns; Model instability: Commercial LLM tools showed temporal variation in performance between initial and repeat testing; Prompt sensitivity: Results may be dependent on specific prompt wording used; Generalizability bias: Findings may not apply to other disciplines or broader physiotherapy field applications; Opaqueness of commercial LLM tools may introduce performance variability; Reference standard (human consensus) may involve decision-making error; In-context learning approach may introduce misleading information; Prompt wording may influence LLM output and limit generalizability; Potential temporal variability in commercial tools (model updates, deployment changes); Reference standard bias: human decision-making error in the reference standard could affect performance metrics; Selection bias: in-context learning may introduce misleading information influencing performance metrics; Model opacity: commercial web-based LLMs are 'black-box' systems with opaque architectures; Temporal variability: significant differences in performance between initial and repeat testing (especially Copilot); Prompt sensitivity: results depend on specific prompt wording which may not generalize; Interrater reliability concern: initial human interrater agreement was moderate (Cohen's kappa 0.48)

Limitations

  • The authors acknowledged several limitations: "First, commercial, web-based LLM tools are opaque at the architectural level, and their performance may not be entirely stable." Additionally, "evaluation metrics depend on the reference standard and were used as in-context learning
  • The reference standard may involve a degree of human decision-making error, and the use of in-context learning may inadvertently mean that misleading information may be provided." Furthermore, "the prompts used are tailored to this proof-of-concept study, and prompt wording may influence LLM output, and thus, the findings may not generalize other disciplines."

Open questions raised

  • Lack of evaluation of LLMs for identifying relevant trials across the broader physiotherapy field beyond musculoskeletal specialty
  • Limited evidence on temporal reproducibility and stability of commercial LLM performance
  • Need for research on using LLMs to perform full-text article screening with richer contextual information
  • Investigation of hybrid screening workflows combining LLM-based and traditional machine learning or NLP techniques as parallel reviewers
  • Use of LLMs to identify duplicates among search results
  • Application of these approaches to other disciplines and evidence synthesis databases
Data: Primary dataset: 10,793 physiotherapy trial records screened from targeted searches across Medline via Ovid, American Psychological Association PsycINFO via Ovid, AMED via Ovid, Embase via Ovid, Cumulative Index of Nursing and Allied Health Literature via EBSCO-host, and Cochrane Central Register of Controlled Trials; Physiotherapy Evidence Database (PEDro; www.pedro.org.au) - a freely available database indexing over 66,000 records; Labeled dataset of 4000 records used for in-context learning; Code and setup details publicly available at https://github.com/Kenn0918/SydneyUniPEDroCode: https://github.com/Kenn0918/SydneyUniPEDro (Details of codes and setup of exploratory approaches to automation); https://github.com/Kenn0918/SydneyUniPEDro - Details of codes and setup for API-based implementation, machine learning, and natural language processing approaches; https://github.com/Kenn0918/SydneyUniPEDroExtracted from: pdfAgreement 51%

Explore related topics

Related papers