Study protocol for evaluating automation of systematic review processes with EPPI-Reviewer and Copilot 365 in updating the cataract evidence gap map
Bhavisha Virendrakumar, Hugh Sharma Waddington, Pauline Scheelbeek, Emma Jolley, Elena Schmidt · Systematic Reviews · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s13643-026-03101-4
Methodology & findings
Study design
Mixed-methods study protocol combining manual and automated screening.
Primary method
The protocol specifies the use of: (1) proportion calculations for relevant references prioritised by prioritisation screening and relevant citations missed; (2) Cohen's Kappa for consistency assessment in full-text screening and automated data extraction/appraisal comparison; (3) time comparison analysis between manual extraction and Copilot 365-assisted extraction.
Main result
This is a study protocol, not a completed empirical study with results. The paper states that "This study will offer insights into ER's accuracy in screening small samples of citations and potentially guide future applications in this context. Additionally, by evaluating Copilot 365, which shares similar features with other AI tools, we will gain a broader understanding of its applicability and limitations in evidence synthesis, making the results relevant to other AI applications in this field." No actual findings are reported as this is a protocol paper describing planned research.
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "This study will offer insights into ER's accuracy in screening small samples of citations and potentially guide future applications in this context. Additionally, by evaluating Copilot 365, which shares similar features with other AI tools, we will gain a broader understanding of its applicability and limitations in evidence synthesis, making the results relevant to other AI applications in this field."
Risk of bias
Selection bias: Limited specification of study selection criteria in protocol; No blinding mentioned for reviewers; Conflict resolution process between two reviewers not fully detailed; Limited sample specification (20-40% thresholds may not be representative); Selection bias in choice of AI tools studied; Potential training bias in AI models; Reviewer bias in manual screening (mitigated by dual independent screening); Limited sample sizes for validation (20-40% manual inspection thresholds); Potential selection bias in the 20% and 40% manual screening thresholds for prioritisation screening evaluation; Potential reviewer bias from the two independent reviewers despite conflict resolution procedures; Unclear generalizability of results from evaluating AI tools on a specific evidence gap map update
Limitations
- The authors note that "the optimal threshold for accuracy remains unclear" for EPPI-Reviewer, and acknowledge that "there is no evidence on the effectiveness of any version of Copilot in systematic review and EGM processes." Additionally, the authors state this is a study protocol, not yet implemented, so no empirical limitations from actual results are available.
Open questions raised
- The protocol identifies that 'Although ER shows promise in speeding up screening, the optimal threshold for accuracy remains unclear.' Additionally, 'there is no evidence on the effectiveness of any version of Copilot in systematic review and EGM processes.' The study aims to address these gaps.
- The authors identify gaps regarding: (1) the optimal threshold for EPPI-Reviewer prioritisation screening accuracy; (2) lack of evidence on Copilot effectiveness in systematic review and EGM processes; (3) applicability of AI tools across different stages of evidence synthesis; and (4) broader understanding needed of AI tool limitations in evidence synthesis.
- The authors identify several research gaps: (1) the optimal threshold for accuracy in EPPI-Reviewer screening remains unclear; (2) there is no evidence on the effectiveness of any version of Copilot in systematic review and evidence gap map processes; (3) limited understanding of the applicability and limitations of AI tools like Copilot 365 in evidence synthesis workflows.
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations