12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automating data extraction in meta-research: A multi-model benchmark in network psychometrics papers

Benjamin Šimsa, artem buts, Ivan Ropovik, Matúš Adamkovič · Behavior Research Methods · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
1
Citations
10.95
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3758/s13428-026-03052-7

Methodology & findings

Study design

Systematic benchmark evaluation comparing five large language models (LLMs) on metascientific data extraction tasks.

Sample

N = 43, 6 groups

Primary method

Generalized linear mixed model (GLMM) with random intercepts for papers and variables; log-odds regression predicting correct extraction (1) vs. incorrect (0); Brier score analysis (mean squared difference between assigned confidence and actual outcome); calibration curves with 10 bins; Area Under the Receiver Operating Characteristic Curve (AUC) for resolution. All analyses conducted in R. API-based extraction with JSON output formatting. Binary and multi-class grading scheme applied to verify accuracy.

Main result

The study found that "Two Anthropic models (Claude 4.6 Opus and Claude 4.5 Haiku) emerged as the two models with the greatest accuracy (91.3% correct responses for Opus, 90.6% for Haiku)." Additionally, "the best-performing models achieved relatively high accuracy for one-shot extraction of a range of metascientific variables from psychological research papers" with "variables such as the title of the paper was extracted with 100% accuracy across all models," while "variables such as the packages used for the estimation (72.6% correct), whether the analytical code was declared to be publicly available (75.4% correct), and the devices used for collecting ESM data (75.5% correct) were found to be more challenging for the models."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/positivist

Author conclusions

The authors conclude that "state-of-the-art LLMs can achieve relatively high one-shot accuracy when extracting diverse metascientific variables from psychological research papers, offering substantial gains in speed and scalability. While not yet a substitute for human judgment, their integration into hybrid workflows can streamline meta-research on an unprecedented scale." They further state that "By providing a reproducible, open framework for automated data extraction-built on API access, structured JSON outputs, and modular and transparent evaluation-our paper aims to offer both the tools and evidence needed to accelerate effective LLM-human integration in solving meta-research questions." They also note that given "median total error rates for human coding are approximately 25% (Mathes et al., 2017), implementing a hybrid pipeline could enhance both the scalability of meta-research workflows and the cross-validation of coding accuracy."

Risk of bias

Reference dataset derived from prior human coding, not treated as error-free ground truth; Single one-shot prompt design without model-specific optimization or few-shot prompting; Each model-paper pair evaluated only once (no replication runs to assess within-model variability); Manual verification procedure for flagged cells may introduce reviewer bias; Limited to psychology subfield with specific reporting conventions; Reference dataset bias: Authors relied on prior human-coded data (Blanchard et al. 2022) as ground truth; 30.8% of manually verified reference cells required modification, indicating potential errors in reference standard; Selection bias: 43 papers from narrow psychology subfield (network psychometrics, ESM/EMA studies); limited to papers published before April 2021; not representative of broader research domains; Single extraction run per model-paper pair: No replication testing within models; unable to assess run-to-run variability; Asymmetric information access: Human coders could access supplementary materials, citations, external databases; LLMs limited to primary document text and pretraining knowledge; Prompt optimization bias: All models received same single-prompt instruction without model-specific optimization or few-shot examples; differences may reflect prompt-model fit rather than intrinsic capability; Temperature setting asymmetry: Claude models set to temperature 0.0 (reproducibility); OpenAI models temperature setting unavailable, potentially introducing variability differences; Grading strictness: Full-list matching requirement (e.g., software packages); any omission or extra item marked incorrect; may penalize partial correctness; Reference dataset bias: Authors note they treated the human-coded reference dataset as 'an operational benchmark derived from prior human coding' rather than error-free ground truth, potentially introducing systematic bias if the original coding contained errors; Single extraction run per model-paper pair: Design does not account for run-to-run variability or stochastic effects in model outputs; Prompt design bias: Single one-shot prompt used for all variables; no comparison against alternative prompt strategies, potentially favoring or disfavoring certain models; Information asymmetry: LLMs limited to text of primary documents while human coders could access supplementary files, citations, and external databases; Narrow domain specificity: Dataset limited to 43 papers from narrow subfield of psychology (network psychometrics with intensive longitudinal data), limiting generalizability; Variable-specific operationalization effects: Authors acknowledge that variable-level performance reflects both extraction difficulty and 'how individual variables were operationalized in the prompt (e.g., the level of contextual detail or specificity provided)' which cannot be independently assessed

Limitations

  • The study states that "The imperfect accuracy observed in this study (75-90%) likely reflects a combination of interacting factors rather than a single limiting source
  • These include properties of the variables (e.g., whether information is explicitly stated versus requiring interpretation), characteristics of the source papers, the design and specificity of the prompt, and model-specific limitations in handling ambiguity or integrating information across the document." Additionally, "human coders in this study had the freedom to consult supplementary files, follow citation trails, or access external databases, whereas LLMs were limited to the text of the primary document and whatever background knowledge they retained from pretraining
  • This asymmetry likely depresses the models' true potential, suggesting that figures reported here should be interpreted as lower-bound estimates." Furthermore, "each model-paper pair was evaluated using a single extraction run," and "The present evaluation was confined to a single dataset of 43 papers drawn from a narrow subfield of psychology with specific reporting conventions, terminology, and article structure."

Open questions raised

  • Sources of error in LLM extraction (variable properties, source paper characteristics, prompt design, model limitations) cannot be fully disentangled in current design
  • Limited comparison of single one-shot prompt against alternative designs (shorter prompts, role-prompt variants, single-question workflows)
  • No exploration of whether prompt engineering aligned with model-specific optimization strategies or few-shot prompting yields improvements
  • Lack of generalizability assessment across disciplines and document types
  • Need for systematic analysis of whether extraction errors cluster in papers with specific reporting styles or structural characteristics
  • Limited exploration of within-model variability due to single extraction run per model-paper pair
Data: Human-coded reference dataset (Blanchard et al. 2022); Study data, code, and materials; Blanchard et al. (2022) reference dataset: https://osf.io/mpj89; Study data, code, and materials: https://osf.io/eqt78/; Meta-research dataset; Study materials and dataCode: OSF (Open Science Framework): https://osf.io/eqt78/; OSF repository with Python code for data extraction, R script for preprocessing/analysis/visualization: https://osf.io/eqt78/; Open Science Framework (OSF): https://osf.io/eqt78/Extracted from: pdfAgreement 50%

Explore related topics

Related papers