Using Elicit AI research assistant for data extraction in systematic reviews: A feasibility study across environmental and life sciences
Malgorzata Lagisz, Ayumi Mizuno, Kyle Morrison, Pietro Pollo, Lorenzo Ricolfi, Yefeng Yang et al. · Research Synthesis Methods · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1017/rsm.2026.10080
Methodology & findings
Study design
Mixed-methods feasibility study involving: (1) prompt development phase using 8 training studies per review to develop and iteratively refine data extraction prompts in Elicit until achieving >87% accuracy (maximum 5 iterations); (2) test phase using 8 new studies per review to evaluate accuracy against human-extracted gold standard data; (3) retest phase using independent Elicit user accounts to assess repeatability; (4) high-accuracy test phase comparing algorithm versions.
Sample
N = 560, 12 groups
Primary method
Analyses conducted in R v.4.5.0. Methods include: Pearson's Chi-squared test for association between variable type and success/failure (Chi-squared = 5.54, df = 3, p = 0.14); Pearson correlation (r = 0.38, t = 3.38, df = 68, p = 0.001) between development and test phase accuracy; Chi-squared test comparing human vs. Elicit error rates (Chi-squared = 11.856, df = 1, p = 0.0006); Chi-squared tests comparing TEST-RETEST (Chi-squared = 0.125, df = 1, p = 0.724) and TEST-HATEST accuracies (Chi-squared = 3.734, df = 1, p = 0.053). Manual comparison of extracted values to gold standard data; no automated comparisons due to need to account for typos, partial matches, and semantic equivalence.
Main result
The study found that "during the prompt development phase, we were unable to reach our extraction accuracy threshold (87%) for 20 out of 90 tested extraction variables within five iterations of prompt refinement" and that "the accuracy of nearly one-third of the variables declined when we applied the same prompts to a new set of studies during the testing phase." Additionally, "almost 90% of the RETEST-extracted values matched exactly the TEST-extracted values (476 out of 536)" indicating high repeatability across user accounts, though "supporting quotes and reasoning provided by Elicit matched in only 46% and 30% of cases, respectively."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist - systematic evaluation of AI tool performance against gold standard human-extracted data
Author conclusions
The authors conclude: "Our study revealed that, in the Elicit platform, (1) developing prompts that reach a predefined level of accuracy in data extractions requires considerable effort, (2) accuracy of extractions varies substantially, (3) data extractions across user accounts are highly repeatable, and (4) distinct LLM algorithms for data extraction produce different results." Further, they recommend: "we recommend integrating Elicit into a modified systematic review workflow for sanity checks or as a secondary extractor alongside a human reviewer for large-scale systematic reviews. A third reviewer could then reconcile discrepancies between human-and Elicit-extracted data, thereby improving efficiency while maintaining high accuracy."
Risk of bias
Small sample size per review (8 studies per phase) limits generalizability; Selection of variables may not represent all types of data extraction challenges; Manual comparison of Elicit outputs with gold standard introduced subjective interpretation of semantically equivalent answers; Excluded studies with unparseable PDF formats (scanned/old documents), creating potential selection bias toward modern, well-formatted PDFs; Human-extracted gold standard data may contain errors (though authors corrected 8 detected errors); Prompt development involved iterative refinement potentially introducing optimization bias; Selection bias in systematic reviews used as basis (though noted they were pre-registered and published); Small sample size per review (8 studies per phase); Subjective manual comparison of Elicit outputs to gold standard (though authors accounted for semantic equivalence); Human error in original gold standard data (detected 8 errors, <1%); Timing of algorithm upgrade during study (TEST vs HATEST comparison affected by Elicit platform changes); Gold standard data from 26 new variables created by only 2 human extractors; Selection bias in variable choice (37 out of 50 variables were specific to single reviews); Small sample size per review (8 studies per phase) limiting generalizability; Subjectivity in manual comparison of extracted data and interpretation of partial matches; Exclusion of studies with unparsable PDFs (scanned/old formats) may bias toward modern, well-formatted articles; Human error in gold standard data creation (though 98.6% accuracy reported); Potential ceiling effects due to researcher involvement in original systematic reviews
Limitations
- The authors state that "the number of studies evaluated per review in each phase was relatively small, which limits the precision of our estimates of success rates and extraction accuracy
- Additionally, we did not attempt to extract numerical values that represent (or that could be used to calculate) effect sizes, as these are often presented in figures or tables, which Elicit currently cannot process." Furthermore, they note "we could not compare extraction times between automated and manual workflows due to a lack of timing data at the variable level for human extractors in the original reviews."
Open questions raised
- The authors identify the need for more detailed documentation of data extractions in published reviews, noting that "metadata (i.e., variable descriptions) from original reviews (conducted exclusively by human reviewers) were often vague, requiring considerable effort to create clear and precise prompts for Elicit. This raises some concerns about the reusability and repeatability of data from published reviews and calls for more detailed documentation of data extractions performed by researchers." They also note that further research is needed on Elicit's performance beyond medical fields.
- The authors identify that "metadata (i.e., variable descriptions) from original reviews (conducted exclusively by human reviewers) were often vague, requiring considerable effort to create clear and precise prompts for Elicit. This raises some concerns about the reusability and repeatability of data from published reviews and calls for more detailed documentation of data extractions performed by researchers." They also note challenges with: hallucinations, misinterpretations, limitations in accessing supplementary materials/figures/tables, missing metadata extraction, and difficulties with older PDF formats.
- Authors identify need for: more detailed documentation of data extractions in published systematic reviews; improved metadata quality from published reviews; further evaluation outside medical fields; investigation of time-cost benefit analysis of Elicit versus manual extraction for large-scale reviews; addressing hallucination, misinterpretation, and overinterpretation issues in LLM-based extraction
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations