12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automated Data Extraction by Large Language Models: Assessing Accuracy in Comparison to Human Experts Using the Example of Visible Learning

Thorben Jansen, Lucas W. Liebenow, Nils-Jonathan Schaller, John Hattie, J. Möller · Educational Psychology Review · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s10648-026-10136-5

Methodology & findings

Study design

A three-phase systematic evaluation comparing LLM-extracted and human expert-extracted data from 156 educational meta-analyses.

Sample

N = 156, 3 groups

Primary method

Percentage agreement, Intraclass Correlation Coefficient ICC(2,1), Pearson correlation (r), Mean Absolute Error (MAE), Bland-Altman plots for assessing systematic bias and limits of agreement. UpSet plots for visualizing agreement patterns across multiple coders. The authors used R for statistical analysis (code available in Supplementary Material). Python was used for LLM API interaction and data processing with PyPDF2.

Main result

The study found that "Across the three models, percentage agreement ranged from 77% to 81%, while the Intraclass Correlation Coefficients (ICC) were excellent, ranging from 0.96 to 0.97" when comparing LLM extractions against the gold standard. Furthermore, "The independent author coding achieved 86% agreement (ICC = 0.95) with the gold standard and 91% agreement (ICC = 0.98) with the silver standard. The LLMs' accuracy fell within a narrow band of this high-quality human performance." This demonstrates that LLMs achieved comparable accuracy to expert human coders in data extraction from educational meta-analyses.

Reports effect sizes.

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude: "The time-consuming nature of data extraction hinders the merits of SOMA. Our study demonstrates that for a specific data extraction task LLMs performed on par with expert humans. Thereby, we provide an empirical foundation for responsible use of LLMs for data extraction." Furthermore, they state: "our results and the emerging literature indicate that a human-LLM hybrid workflow is superior to a human-only approach under identical resource constraints." They also conclude that "LLMs can already be valuable for extracting effect sizes, the number of included studies, and the number of included participants from meta-analyses" based on findings that "LLMs operating at comparable levels of accuracy can be leveraged to substantially accelerate evidence synthesis without increasing overall error rates."

Risk of bias

Confirmation bias in adjudication process—adjudicators were study authors involved in initial coding; Overfitting risk from iterative prompt engineering on pilot set of 10 meta-analyses; Potential train/test contamination—Visible Learning database likely included in LLM training data; Selection bias—human coders were experts, not representative of typical review teams with mixed experience levels; Publication bias in source meta-analyses—not assessed; Non-blinded adjudication procedure during silver and gold standard creation; Selection bias: Specialized expert coders not representative of typical review teams; Confirmation bias: Adjudicators who conducted initial coding also performed adjudication; Overfitting: Intensive prompt engineering on pilot sample of 10 meta-analyses; Context bias: Optimal extraction conditions with clearly reported numeric variables; Training data contamination: Visible Learning database likely included in LLM training data; Attrition bias: 44 of 200 randomly selected meta-analyses excluded (23% exclusion rate); Confirmation bias in adjudication process (conducted by study authors involved in initial coding); Not fully blinded adjudication procedure; Potential train/test contamination risk with LLMs (Visible Learning database likely in training data); Overfitting of prompts to pilot dataset (n=10); Selection bias: expert human coders only (not representative of typical review teams with mixed experience); Assumption that unanimous coder agreement indicates correctness without independent verification

Limitations

  • "Our study deliberately focused on three standardized numeric variables (effect size, number of studies, and number of participants) extracted from educational meta-analyses in the Visible Learning database
  • This is a comparatively well-structured extraction setting because these variables are central to meta-analytic reporting and are often repeated across tables, abstracts, and results sections
  • Accordingly, our accuracy estimates should be interpreted as upper bounds (best cases) for both humans and LLMs and should not be generalized to extraction tasks that require construct interpretation, judgment, or classification." Additionally, "the adjudication process was conducted by the study authors, who were also involved in the initial author coding
  • While it is standard practice in meta-analytic research for the author team to resolve discrepancies, this design is not fully blinded and introduces a potential for confirmation bias, where adjudicators might unintentionally favor their original ratings."

Open questions raised

  • Future research should: (1) evaluate LLMs on less standardized, interpretive coding tasks requiring judgment and construct interpretation; (2) determine optimal number of coders (human or LLM) needed for data validation; (3) assess open-source LLM alternatives; (4) examine performance with less-experienced human coders; (5) conduct prospective replications using novel, non-public datasets to eliminate train/test contamination risk; (6) evaluate responsible integration patterns (LLM as quality assurance tool, extraction assistant, or second independent coder); (7) investigate performance across different meta-analytic reporting structures and domains beyond educational achievement.
  • Generalizability to extraction tasks requiring construct interpretation, judgment, or classification (beyond numeric variables)
  • Performance with novice or less-experienced coders
  • How many independent coders (human or LLM) are needed to validate data: "For future studies, it would be especially interesting to determine how many human coders or LLMs, when extracting data across multiple runs or with different prompts, need to agree to ensure the data is validated."
  • Performance with open-source alternative LLMs
  • Long-term reproducibility and sustainability of proprietary LLM-based approaches
Data: All data, LLM prompts, and analysis scripts available at https://osf.io/q5n7f (Open Science Framework). The study extracted data from 156 educational meta-analyses randomly selected from the Visible Learning database (https://www.visiblelearningmetax.com).; All data and analysis scripts available at: https://osf.io/q5n7f (OSF - Open Science Framework); All data, LLM prompts, and analysis scripts available at: https://osf.io/q5n7f; 156 meta-analyses randomly selected from Visible Learning database (https://www.visiblelearningmetax.com)Code: R analysis scripts available in Supplementary Material; Python application code for LLM-based data extraction available in Supplementary Material C; Python application code for LLM data extraction available in Supplementary Material C; R code for statistical analysis available in Supplementary Material; Python application code for LLM-based data extraction available in Supplementary Material C at https://osf.io/q5n7f; Analysis scripts (R code) available in Supplementary Material at https://osf.io/q5n7fExtracted from: pdfAgreement 49%

Explore related topics

Related papers