Active Learning In Systematic Literature Screening: Benchmarking Ai-Assisted Screening Against Human Consensus
Mickael Ringeval, Guy Paré, Grégory Vial, Aude Motulsky · Journal of the Association for Information Systems · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Exploratory benchmarking design comparing active learning (AL) assisted screening using two off-the-shelf tools (ASReview and Covidence) against validated human consensus screening decisions.
Main result
The study found that "ASReview achieved a recall of 100%, as it successfully identified all 28 relevant studies included in the final sample" and "screening only 746 of the 5,808 references corresponds to a workload reduction of 87.2% compared to manual screening, indicating substantial efficiency gains without loss of comprehensiveness." In contrast, "Covidence achieved a recall of 17.8%, as it identified 5 of the 28 relevant studies included in the final sample" when using the same data-driven stopping rule, though extending the process to full recall required "167 minutes and 1,279 screened references."
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that "This exploratory study advances our understanding of AL tools in the screening phase of literature reviews. By comparing the performance of ASReview and Covidence with traditional manual screening, the findings illustrate how AI-assisted approaches can enhance review processes, particularly by improving efficiency under appropriate calibration and stopping conditions." They further state that "Although our findings confirm that ASReview performs strongly, it should not yet be viewed as a standalone replacement for manual screening. This study was exploratory, designed to assess the potential rather than the full maturity of AL tools in systematic reviews. At present, ASReview is best used as a complementary or validation tool alongside manual screening."
Risk of bias
Single researcher conducting screening with both tools (potential for consistency/learning effects between tool applications); Temporal gap between ASReview screening (2024) and Covidence screening (one month later) which could affect recall due to learning; Tool-specific algorithmic opacity limits ability to assess whether systematic biases exist in prioritization; Calibration procedures differ substantially between tools, potentially introducing systematic bias in comparisons; No blinding mentioned for researcher performing screening; Single reviewer conducted AL screening (no inter-rater reliability assessment); Potential selection bias in calibration choices for ASReview (using one reference from final sample); Lack of transparency in proprietary algorithms (Covidence opaque decision logic); Temporal separation between ASReview and Covidence screening (one month apart); Limited dataset (5,808 references) - authors note AL tools perform better on larger datasets; Absence of standardized calibration procedures across tools; Single researcher conducting AL screening (no inter-rater reliability checks); Potential researcher familiarity bias with ASReview (conducted first, followed by Covidence after one month delay); Possible differential calibration effects between tools due to different requirements and user guidance availability; Manual baseline screening conducted by four reviewers with consensus procedure, while AL screening by single researcher; Lack of blinding regarding reference standard when conducting AL screening; Tool-specific algorithmic biases not independently assessed
Limitations
- The authors state that "AL tools typically perform better and are more beneficial for larger datasets, where manual screening becomes increasingly demanding in terms of time and effort" and that "Our study involved 5,808 references, which represents a dataset size commonly encountered in systematic reviews." Additionally, "calibration procedures play a critical role in determining the effectiveness of AL-assisted screening
- However, guidance on how to perform calibration remains limited and is often tool-specific rather than aligned with the methodological objectives of the review." They also note that "AL tools are typically implemented as digital platforms that evolve over time, and researchers have limited control over updates or changes to the underlying models," limiting reproducibility.
Open questions raised
- Absence of standardized evaluation frameworks and consistent performance metrics for assessing AI-supported screening in systematic reviews
- Limited empirical evidence on whether active learning-assisted screening can reliably preserve relevant studies while achieving meaningful efficiency gains
- Lack of guidance on how to perform calibration; guidance remains limited and is often tool-specific rather than aligned with methodological objectives
- Need for clearer methodological guidance on stopping rules; future research could compare existing stopping heuristics to help researchers select appropriate methods
- Transparency and interpretability challenges in many AL systems regarding algorithms and decision rationale
- Need for curriculum design, workshop development, and training programs to equip researchers with knowledge about AL tools
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations