Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning
Thorben Jansen, Lucas W. Liebenow, Nils-Jonathan Schaller, John Hattie, J. Möller · Educational Psychology Review · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s10648-026-10136-5
Methodology & findings
Study design
A three-phase systematic evaluation comparing LLM-extracted and human expert-extracted data from 156 educational meta-analyses.
Sample
N = 156, 3 groups
Primary method
Percentage agreement, Intraclass Correlation Coefficient ICC(2,1), Pearson correlation (r), Mean Absolute Error (MAE), Bland-Altman plots for assessing systematic bias and limits of agreement. UpSet plots for visualizing agreement patterns across multiple coders. The authors used R for statistical analysis (code available in Supplementary Material). Python was used for LLM API interaction and data processing with PyPDF2.
Main result
The study found that "Across the three models, percentage agreement ranged from 77% to 81%, while the Intraclass Correlation Coefficients (ICC) were excellent, ranging from 0.96 to 0.97" when comparing LLM extractions against the gold standard. Furthermore, "The independent author coding achieved 86% agreement (ICC = 0.95) with the gold standard and 91% agreement (ICC = 0.98) with the silver standard. The LLMs' accuracy fell within a narrow band of this high-quality human performance." This demonstrates that LLMs achieved comparable accuracy to expert human coders in data extraction from educational meta-analyses.
Reports effect sizes.
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude: "The time-consuming nature of data extraction hinders the merits of SOMA. Our study demonstrates that for a specific data extraction task LLMs performed on par with expert humans. Thereby, we provide an empirical foundation for responsible use of LLMs for data extraction." Furthermore, they state: "our results and the emerging literature indicate that a human-LLM hybrid workflow is superior to a human-only approach under identical resource constraints." They also conclude that "LLMs can already be valuable for extracting effect sizes, the number of included studies, and the number of included participants from meta-analyses" based on findings that "LLMs operating at comparable levels of accuracy can be leveraged to substantially accelerate evidence synthesis without increasing overall error rates."
Risk of bias
Confirmation bias in adjudication process—adjudicators were study authors involved in initial coding; Overfitting risk from iterative prompt engineering on pilot set of 10 meta-analyses; Potential train/test contamination—Visible Learning database likely included in LLM training data; Selection bias—human coders were experts, not representative of typical review teams with mixed experience levels; Publication bias in source meta-analyses—not assessed; Non-blinded adjudication procedure during silver and gold standard creation; Selection bias: Specialized expert coders not representative of typical review teams; Confirmation bias: Adjudicators who conducted initial coding also performed adjudication; Overfitting: Intensive prompt engineering on pilot sample of 10 meta-analyses; Context bias: Optimal extraction conditions with clearly reported numeric variables; Training data contamination: Visible Learning database likely included in LLM training data; Attrition bias: 44 of 200 randomly selected meta-analyses excluded (23% exclusion rate); Confirmation bias in adjudication process (conducted by study authors involved in initial coding); Not fully blinded adjudication procedure; Potential train/test contamination risk with LLMs (Visible Learning database likely in training data); Overfitting of prompts to pilot dataset (n=10); Selection bias: expert human coders only (not representative of typical review teams with mixed experience); Assumption that unanimous coder agreement indicates correctness without independent verification
Limitations
- "Our study deliberately focused on three standardized numeric variables (effect size, number of studies, and number of participants) extracted from educational meta-analyses in the Visible Learning database
- This is a comparatively well-structured extraction setting because these variables are central to meta-analytic reporting and are often repeated across tables, abstracts, and results sections
- Accordingly, our accuracy estimates should be interpreted as upper bounds (best cases) for both humans and LLMs and should not be generalized to extraction tasks that require construct interpretation, judgment, or classification." Additionally, "the adjudication process was conducted by the study authors, who were also involved in the initial author coding
- While it is standard practice in meta-analytic research for the author team to resolve discrepancies, this design is not fully blinded and introduces a potential for confirmation bias, where adjudicators might unintentionally favor their original ratings."
Open questions raised
- Future research should: (1) evaluate LLMs on less standardized, interpretive coding tasks requiring judgment and construct interpretation; (2) determine optimal number of coders (human or LLM) needed for data validation; (3) assess open-source LLM alternatives; (4) examine performance with less-experienced human coders; (5) conduct prospective replications using novel, non-public datasets to eliminate train/test contamination risk; (6) evaluate responsible integration patterns (LLM as quality assurance tool, extraction assistant, or second independent coder); (7) investigate performance across different meta-analytic reporting structures and domains beyond educational achievement.
- Generalizability to extraction tasks requiring construct interpretation, judgment, or classification (beyond numeric variables)
- Performance with novice or less-experienced coders
- How many independent coders (human or LLM) are needed to validate data: "For future studies, it would be especially interesting to determine how many human coders or LLMs, when extracting data across multiple runs or with different prompts, need to agree to ensure the data is validated."
- Performance with open-source alternative LLMs
- Long-term reproducibility and sustainability of proprietary LLM-based approaches
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations