12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Active Learning In Systematic Literature Screening: Benchmarking Ai-Assisted Screening Against Human Consensus

Mickael Ringeval, Guy Paré, Grégory Vial, Aude Motulsky · Journal of the Association for Information Systems · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Exploratory benchmarking design comparing active learning (AL) assisted screening using two off-the-shelf tools (ASReview and Covidence) against validated human consensus screening decisions.

Main result

The study found that "ASReview achieved a recall of 100%, as it successfully identified all 28 relevant studies included in the final sample" and "screening only 746 of the 5,808 references corresponds to a workload reduction of 87.2% compared to manual screening, indicating substantial efficiency gains without loss of comprehensiveness." In contrast, "Covidence achieved a recall of 17.8%, as it identified 5 of the 28 relevant studies included in the final sample" when using the same data-driven stopping rule, though extending the process to full recall required "167 minutes and 1,279 screened references."

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "This exploratory study advances our understanding of AL tools in the screening phase of literature reviews. By comparing the performance of ASReview and Covidence with traditional manual screening, the findings illustrate how AI-assisted approaches can enhance review processes, particularly by improving efficiency under appropriate calibration and stopping conditions." They further state that "Although our findings confirm that ASReview performs strongly, it should not yet be viewed as a standalone replacement for manual screening. This study was exploratory, designed to assess the potential rather than the full maturity of AL tools in systematic reviews. At present, ASReview is best used as a complementary or validation tool alongside manual screening."

Risk of bias

Single researcher conducting screening with both tools (potential for consistency/learning effects between tool applications); Temporal gap between ASReview screening (2024) and Covidence screening (one month later) which could affect recall due to learning; Tool-specific algorithmic opacity limits ability to assess whether systematic biases exist in prioritization; Calibration procedures differ substantially between tools, potentially introducing systematic bias in comparisons; No blinding mentioned for researcher performing screening; Single reviewer conducted AL screening (no inter-rater reliability assessment); Potential selection bias in calibration choices for ASReview (using one reference from final sample); Lack of transparency in proprietary algorithms (Covidence opaque decision logic); Temporal separation between ASReview and Covidence screening (one month apart); Limited dataset (5,808 references) - authors note AL tools perform better on larger datasets; Absence of standardized calibration procedures across tools; Single researcher conducting AL screening (no inter-rater reliability checks); Potential researcher familiarity bias with ASReview (conducted first, followed by Covidence after one month delay); Possible differential calibration effects between tools due to different requirements and user guidance availability; Manual baseline screening conducted by four reviewers with consensus procedure, while AL screening by single researcher; Lack of blinding regarding reference standard when conducting AL screening; Tool-specific algorithmic biases not independently assessed

Limitations

  • The authors state that "AL tools typically perform better and are more beneficial for larger datasets, where manual screening becomes increasingly demanding in terms of time and effort" and that "Our study involved 5,808 references, which represents a dataset size commonly encountered in systematic reviews." Additionally, "calibration procedures play a critical role in determining the effectiveness of AL-assisted screening
  • However, guidance on how to perform calibration remains limited and is often tool-specific rather than aligned with the methodological objectives of the review." They also note that "AL tools are typically implemented as digital platforms that evolve over time, and researchers have limited control over updates or changes to the underlying models," limiting reproducibility.

Open questions raised

  • Absence of standardized evaluation frameworks and consistent performance metrics for assessing AI-supported screening in systematic reviews
  • Limited empirical evidence on whether active learning-assisted screening can reliably preserve relevant studies while achieving meaningful efficiency gains
  • Lack of guidance on how to perform calibration; guidance remains limited and is often tool-specific rather than aligned with methodological objectives
  • Need for clearer methodological guidance on stopping rules; future research could compare existing stopping heuristics to help researchers select appropriate methods
  • Transparency and interpretability challenges in many AL systems regarding algorithms and decision rationale
  • Need for curriculum design, workshop development, and training programs to equip researchers with knowledge about AL tools
Data: The screening corpus is derived from a published systematic review (Vial et al., 2025) that was peer-reviewed and published in the Journal of the American Medical Informatics Association. The dataset consists of 5,808 unique references from Scopus and Web of Science databases related to large-scale Electronic Health Record implementations. However, no explicit statement is made regarding public availability of the dataset.; The dataset comprises bibliographic records from Scopus and Web of Science databases related to large-scale Electronic Health Record (EHR) implementations. The dataset includes 5,808 unique references after duplicate removal and corresponds to validated screening decisions from a published review article (Vial et al., 2025) in the Journal of the American Medical Informatics Association. The dataset is not explicitly stated as publicly available in the paper.; EHR implementation systematic review datasetExtracted from: pdfAgreement 52%

Explore related topics

Related papers