12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Study protocol for evaluating automation of systematic review processes with EPPI-Reviewer and Copilot 365 in updating the cataract evidence gap map

Bhavisha Virendrakumar, Hugh Sharma Waddington, Pauline Scheelbeek, Emma Jolley, Elena Schmidt · Systematic Reviews · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s13643-026-03101-4

Methodology & findings

Study design

Mixed-methods study protocol combining manual and automated screening.

Primary method

The protocol specifies the use of: (1) proportion calculations for relevant references prioritised by prioritisation screening and relevant citations missed; (2) Cohen's Kappa for consistency assessment in full-text screening and automated data extraction/appraisal comparison; (3) time comparison analysis between manual extraction and Copilot 365-assisted extraction.

Main result

This is a study protocol, not a completed empirical study with results. The paper states that "This study will offer insights into ER's accuracy in screening small samples of citations and potentially guide future applications in this context. Additionally, by evaluating Copilot 365, which shares similar features with other AI tools, we will gain a broader understanding of its applicability and limitations in evidence synthesis, making the results relevant to other AI applications in this field." No actual findings are reported as this is a protocol paper describing planned research.

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude that "This study will offer insights into ER's accuracy in screening small samples of citations and potentially guide future applications in this context. Additionally, by evaluating Copilot 365, which shares similar features with other AI tools, we will gain a broader understanding of its applicability and limitations in evidence synthesis, making the results relevant to other AI applications in this field."

Risk of bias

Selection bias: Limited specification of study selection criteria in protocol; No blinding mentioned for reviewers; Conflict resolution process between two reviewers not fully detailed; Limited sample specification (20-40% thresholds may not be representative); Selection bias in choice of AI tools studied; Potential training bias in AI models; Reviewer bias in manual screening (mitigated by dual independent screening); Limited sample sizes for validation (20-40% manual inspection thresholds); Potential selection bias in the 20% and 40% manual screening thresholds for prioritisation screening evaluation; Potential reviewer bias from the two independent reviewers despite conflict resolution procedures; Unclear generalizability of results from evaluating AI tools on a specific evidence gap map update

Limitations

  • The authors note that "the optimal threshold for accuracy remains unclear" for EPPI-Reviewer, and acknowledge that "there is no evidence on the effectiveness of any version of Copilot in systematic review and EGM processes." Additionally, the authors state this is a study protocol, not yet implemented, so no empirical limitations from actual results are available.

Open questions raised

  • The protocol identifies that 'Although ER shows promise in speeding up screening, the optimal threshold for accuracy remains unclear.' Additionally, 'there is no evidence on the effectiveness of any version of Copilot in systematic review and EGM processes.' The study aims to address these gaps.
  • The authors identify gaps regarding: (1) the optimal threshold for EPPI-Reviewer prioritisation screening accuracy; (2) lack of evidence on Copilot effectiveness in systematic review and EGM processes; (3) applicability of AI tools across different stages of evidence synthesis; and (4) broader understanding needed of AI tool limitations in evidence synthesis.
  • The authors identify several research gaps: (1) the optimal threshold for accuracy in EPPI-Reviewer screening remains unclear; (2) there is no evidence on the effectiveness of any version of Copilot in systematic review and evidence gap map processes; (3) limited understanding of the applicability and limitations of AI tools like Copilot 365 in evidence synthesis workflows.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 60%

Explore related topics

Related papers