12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Accuracy and efficiency of using artificial intelligence for data extraction in systematic reviews. A noninferiority study within reviews

Kate M O'Brien, Justin Presseau, Serene Yoong, Christophe Lecathelinais, Luke Wolfenden, James A. Thomas et al. · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.02.25.26347053

Methodology & findings

Study design

Noninferiority study within a systematic review (SWAR); two-arm parallel-group design comparing AI-assisted data extraction (using Elicit®) with human-only data extraction across 50 randomized controlled trials.

Sample

N = 50, 2 groups

Primary method

Paired t-tests for comparing accuracy and time-to-complete between arms. Subgroup analysis conducted using t-tests where more than one variable was available. Normality assessed through visual inspection of histograms. Noninferiority assessed by comparing if differences between both arms were within the predefined 10% noninferiority margin and within upper and lower 95% confidence intervals. Significance level α=0.025 (one-tailed for noninferiority) used. Descriptive statistics for participant characteristics, costs, and error types/severity. Analysis conducted in SAS version 9.4 (for accuracy) and Microsoft Excel Version 2504 (for time-to-complete). Sensitivity analysis conducted excluding preparation time.

Main result

The study found that "AI-assisted data extraction using Elicit® showed noninferior accuracy, faster completion times, similar error types and severity, and lower costs compared to human-only extraction." The mean overall accuracy score for the AI-assisted and human-only data extraction arms were 85.8 (range 74-99) and 85.3 (range 72-94) respectively, with a mean difference of 0.57 (95% CI -1.29, 2.43). "AI-assisted data extraction was significantly faster (MD 24.82 mins, 95% CI 18.80, 30.84)."

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude: "The use of Elicit® is noninferior to human-only data extraction in both accuracy and time-to-complete. Both AI-assisted and human-only data extraction made similar types and severity of errors, and AI-assisted data extraction cost less than human-only data extraction." They further state that "AI-assisted data extraction using Elicit® showed noninferior accuracy, faster completion times, similar error types and severity, and lower costs compared to human-only extraction. These efficiency gains, without loss in accuracy suggest AI-assisted data extraction can replace one human-only data extractor in future systematic reviews of RCTs."

Risk of bias

Data contamination risk: Elicit® may have been trained on the same openly published review data used in this trial; Participant blinding: Participants were not blinded to group allocation; only the accuracy assessor (KO) and data analyst were blinded; Limited participant experience: Research assistants had relatively limited experience in systematic review data extraction; Single data extraction without consolidation: No consensus or third-person review process applied, which is atypical for systematic reviews; Unblinded lead author: The lead author (DL) was not blinded to allocation sequence; Data contamination: Elicit® may have been trained on the same publicly available data used in the trial; Limited blinding: Participants were aware of their group allocation (full blinding not possible); Limited experience of extractors: Research assistants had relatively limited experience in systematic review data extraction; Single blinded assessor: Only one assessor (KO) performed accuracy assessment, though assessor was blinded to study arms; Unstructured outcome data: Requiring additional processing before analysis may introduce additional errors; Data contamination risk: Elicit® may have been trained on the same published review and trials used in this study; Limited blinding: Participants knew their group allocation; full blinding was not possible; Benchmark bias mitigated: Use of blinded assessor with predefined rubric against original papers rather than imperfect human-only extraction as reference standard; Selection bias: Only school-based RCTs written in English with readable PDFs processable by Elicit®; Attrition: None reported (2 participants completed all tasks); Confounders: Participant experience levels differed (one had led a systematic review, one had not); participants had different mean time-to-complete (15.8 minutes difference)

Limitations

  • The authors state several key limitations: "First, reliance on an openly published review and publicly available trials as a data source, introduces a potential for data contamination, as Elicit® may have been trained on this same data." Second, "our findings are limited to the version of Elicit® tested, with rapid advancements in LLM and newer features already available (e.g
  • all columns are now high-accuracy), future iterations may yield different and potentially improved results." Third, "outcome data from both arms were unstructured and would require additional processing before analysis, which is an atypical workflow for a well-designed systematic review." Finally, "this trial involved research assistants with relatively limited experience in systematic review data extraction, which may have influenced overall accuracy across both arms."

Open questions raised

  • Need for exploration of different models of AI data extraction, such as two AI-assisted extractors or AI-only extractor with human-only extractor comparison
  • Comparison of AI-assisted to AI-only extraction
  • Integration of Elicit® into living systematic reviews to evaluate cost-effectiveness over multiple updates and mitigate data contamination risk
  • Investigation of efficacy using two AI-assisted arms in parallel or AI tool checking output before human oversight
  • Trials exploring hybrid model approaches (e.g., using AI to assign difficult articles to humans) to replace human extractors
  • Research with more experienced extractors (e.g., five years post-doctoral experience) versus the less experienced extractors used in this trial
Data: Source data from 2022 Cochrane systematic review on obesity prevention interventions (reference 40); 50 school-based RCTs sourced from a 2022 Cochrane systematic review on obesity prevention interventions in children aged 6-18 years (ref 40)Code: Not mentionedExtracted from: pdfAgreement 54%

Explore related topics

Related papers