Evaluating machine learning tools to assist title and abstract screening in systematic literature reviews: a report based on the EULAR RA Management Recommendations Task Force
Victoria Konzett, Faidra Laskou, J S Smolen, Christopher J Edwards, Daniel Aletaha, Désirée van der Heijde et al. · Annals of the Rheumatic Diseases · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1016/j.ard.2026.05.006
Methodology & findings
Study design
Systematic replication study: Four previously conducted systematic literature reviews (SLRs) for EULAR RA management recommendations were manually screened using Rayyan as reference standard.
Sample
N = 16403, 20 groups
Primary method
Descriptive statistics (means, standard deviations, proportions, percentages). Performance metrics evaluated include: recall (true positive rate - sensitivity), precision (positive predictive value), workload reduction (proportion of records screened), and efficiency measurements. Recall curves were used to assess model performance. No inferential statistical tests are explicitly reported.
Main result
The study found that ML-assisted tools achieved significant workload reduction while maintaining reasonable recall. "The average proportion of records that had to be screened ML-assisted tools, compared with manual search, was 22.2% ± 12.8%" with ASReview achieving 14-20% of total abstracts screened per SLR (80-86% workload reduction), SWIFT-Active Screener achieving 40-69% workload reduction, and Research Screener achieving 82-96% workload reduction. However, "between 0 and 3 (4.2%) of records included with the traditional method were missed by the ML tools."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "ML models designed to support title and abstract screening for SLRs hold promise as efficient and reliable tools for generating comprehensive literature summaries, while substantially reducing workload and improving feasibility in a rapidly growing scientific landscape. However, safeguards, such as consistent training, evaluation, and calibration of the models, as well as user education, are essential to prevent bias and low-quality output." Additionally, they state: "In summary, we provide an overview and evaluation of currently available ML models for title and abstract screening of SLRs. Our findings demonstrate good functionality and reliability across different models, supporting the integration of AI into SLR workflows, while considering potential pitfalls, associated risks, and the need for ongoing methodological development and researcher education in a rapidly evolving field."
Risk of bias
Knowledge of manual screening results before conducting ML screening (information bias); Single reviewer conducting semiautomated screening (potential for introduction of reviewer bias); Nonsystematic selection of ML models may introduce selection bias; Lack of standardized time measurement for manual screening (performance bias); Domain-specific focus on rheumatoid arthritis may introduce external validity concerns; Knowledge of manual screening results before AI screening (post-hoc assessment bias); Single reviewer conducting semiautomated screening (no independent verification); Non-systematic selection of ML models evaluated; Domain-specific bias: limited to rheumatology/RA research; Lack of prospective validation design; Potential human error in manual screening not quantified as baseline error rate; Knowledge of manual search results during ML screening (not blinded); Single researcher conducting semiautomated screening; Potential human error amplification by ML algorithms when learning from reviewer decisions; Lack of prospective study design
Limitations
- The authors identified several key limitations: "AI-based screenings were performed after manual searches in this study, meaning that the manual screening results were known to the reviewers
- This is an important limitation of the current work and could obscure a potential challenge when working with active learning algorithms, especially for inexperienced researchers." Additional limitations include "the conduction of semiautomated SLRs by only 1 researcher, the nonsystematic selection of eligible ML models" and "the lack of an exact assessment of time savings compared with the manual search, as manual screening time was not recorded during the original searches." The authors also note "the lack of generalisability across multiple domains needs to be noted, as the current project covered 4 SLRs from a very specific research domain."
Open questions raised
- The authors identify several gaps: (1) Need for validated stopping rules and acceptable performance thresholds for ML-assisted SLRs which currently are not defined; (2) Lack of prospective validation studies using heterogeneous record sets; (3) Need for continuous researcher education and training to prevent uncritical use of AI; (4) Requirement for standardized quality assessment and validation of ML models prior to implementation; (5) Need for scenario-specific recall thresholds and stringency levels; (6) Generalizability testing across multiple research domains beyond rheumatology.
- Validated stopping rules needed for different SLR scenarios (RCTs only vs. those including observational data)
- Lack of defined acceptable performance thresholds for ML-assisted SLRs
- Need for prospective validation studies using heterogeneous record sets
- Replication of ML-based searches on large datasets across multiple domains needed
- Continuous researcher education and training to prevent uncritical use of AI
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations