12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating machine learning tools to assist title and abstract screening in systematic literature reviews: a report based on the EULAR RA Management Recommendations Task Force

Victoria Konzett, Faidra Laskou, J S Smolen, Christopher J Edwards, Daniel Aletaha, Désirée van der Heijde et al. · Annals of the Rheumatic Diseases · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.ard.2026.05.006

Methodology & findings

Study design

Systematic replication study: Four previously conducted systematic literature reviews (SLRs) for EULAR RA management recommendations were manually screened using Rayyan as reference standard.

Sample

N = 16403, 20 groups

Primary method

Descriptive statistics (means, standard deviations, proportions, percentages). Performance metrics evaluated include: recall (true positive rate - sensitivity), precision (positive predictive value), workload reduction (proportion of records screened), and efficiency measurements. Recall curves were used to assess model performance. No inferential statistical tests are explicitly reported.

Main result

The study found that ML-assisted tools achieved significant workload reduction while maintaining reasonable recall. "The average proportion of records that had to be screened ML-assisted tools, compared with manual search, was 22.2% ± 12.8%" with ASReview achieving 14-20% of total abstracts screened per SLR (80-86% workload reduction), SWIFT-Active Screener achieving 40-69% workload reduction, and Research Screener achieving 82-96% workload reduction. However, "between 0 and 3 (4.2%) of records included with the traditional method were missed by the ML tools."

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "ML models designed to support title and abstract screening for SLRs hold promise as efficient and reliable tools for generating comprehensive literature summaries, while substantially reducing workload and improving feasibility in a rapidly growing scientific landscape. However, safeguards, such as consistent training, evaluation, and calibration of the models, as well as user education, are essential to prevent bias and low-quality output." Additionally, they state: "In summary, we provide an overview and evaluation of currently available ML models for title and abstract screening of SLRs. Our findings demonstrate good functionality and reliability across different models, supporting the integration of AI into SLR workflows, while considering potential pitfalls, associated risks, and the need for ongoing methodological development and researcher education in a rapidly evolving field."

Risk of bias

Knowledge of manual screening results before conducting ML screening (information bias); Single reviewer conducting semiautomated screening (potential for introduction of reviewer bias); Nonsystematic selection of ML models may introduce selection bias; Lack of standardized time measurement for manual screening (performance bias); Domain-specific focus on rheumatoid arthritis may introduce external validity concerns; Knowledge of manual screening results before AI screening (post-hoc assessment bias); Single reviewer conducting semiautomated screening (no independent verification); Non-systematic selection of ML models evaluated; Domain-specific bias: limited to rheumatology/RA research; Lack of prospective validation design; Potential human error in manual screening not quantified as baseline error rate; Knowledge of manual search results during ML screening (not blinded); Single researcher conducting semiautomated screening; Potential human error amplification by ML algorithms when learning from reviewer decisions; Lack of prospective study design

Limitations

  • The authors identified several key limitations: "AI-based screenings were performed after manual searches in this study, meaning that the manual screening results were known to the reviewers
  • This is an important limitation of the current work and could obscure a potential challenge when working with active learning algorithms, especially for inexperienced researchers." Additional limitations include "the conduction of semiautomated SLRs by only 1 researcher, the nonsystematic selection of eligible ML models" and "the lack of an exact assessment of time savings compared with the manual search, as manual screening time was not recorded during the original searches." The authors also note "the lack of generalisability across multiple domains needs to be noted, as the current project covered 4 SLRs from a very specific research domain."

Open questions raised

  • The authors identify several gaps: (1) Need for validated stopping rules and acceptable performance thresholds for ML-assisted SLRs which currently are not defined; (2) Lack of prospective validation studies using heterogeneous record sets; (3) Need for continuous researcher education and training to prevent uncritical use of AI; (4) Requirement for standardized quality assessment and validation of ML models prior to implementation; (5) Need for scenario-specific recall thresholds and stringency levels; (6) Generalizability testing across multiple research domains beyond rheumatology.
  • Validated stopping rules needed for different SLR scenarios (RCTs only vs. those including observational data)
  • Lack of defined acceptable performance thresholds for ML-assisted SLRs
  • Need for prospective validation studies using heterogeneous record sets
  • Replication of ML-based searches on large datasets across multiple domains needed
  • Continuous researcher education and training to prevent uncritical use of AI
Data: The datasets used were from 4 SLRs conducted to inform the 2025 EULAR recommendations for rheumatoid arthritis management. Original SLR publications are referenced [16, 17] but specific dataset accessibility is not explicitly stated in this paper.Code: ASReview is noted as open-source (https://asreview.nl) with Apache 2.0 license. Other tools mentioned are proprietary or cloud-based.Extracted from: pdfAgreement 54%

Explore related topics

Related papers