12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Automated reproducibility assessments in the social and behavioral sciences using large language models

Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational simulation study using an agentic large language model (LLM) pipeline.

Main result

The study found that "LLMs can reproduce a substantial share of published findings in the social and behavioral sciences" with "the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a ±0.05 tolerance in Cohen's d) in 24% of studies." In the subset with human reanalyses, "the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a ±0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%)"

Research paradigm

Empirical/quantitative positivism with computational verification

Author conclusions

The authors conclude: "In sum, our findings suggest that LLMs can offer a path toward more scalable reproducibility assessment by automatically performing systematic reanalyses that are otherwise costly to conduct manually. Given the current capabilities and limitations of LLMs, such automated reanalysis should be viewed as a scalable screening tool rather than a substitute for expert judgment. Nevertheless, it could support the research community in quality control and provide a first-pass check for scientific journals, which may ultimately help improve rigor and reproducibility in empirical research." They further note: "Rather than replacing expert judgment, LLM-based reanalysis could support editors, reviewers, and meta-scientists by providing a scalable first-pass check of whether reported findings can be recovered from the available data and study materials."

Risk of bias

Training data contamination: LLMs may have encountered study materials during pretraining; Memorization risk: Model may implicitly leverage prior exposure to studies; Model bias toward original reported results: LLM estimates correlated more strongly with original findings (r=0.46) than with human reanalyses (r=0.11); Selection bias: Corpus limited to studies with available data and extractable claims; Measurement bias: Cohen's d conversions assume standardized effect size comparability across heterogeneous analytical settings; Training data contamination: LLMs may have encountered published studies during pretraining, potentially inflating reproducibility rates; Selection bias: Corpus limited to studies from psychology, economics, and political science with available data; Anchoring bias: LLM pipeline may be biased toward recovering originally reported results rather than exploring alternative analytical paths; Prompt sensitivity: Different prompt framings could influence model behavior (though sensitivity analysis showed only 4 percentage point differences); Methodological assumptions: Conversion of heterogeneous effect sizes to Cohen's d rests on assumptions not universally applicable; Training data contamination: LLMs may have been exposed to studies during pretraining, potentially inflating reproducibility rates; Selection bias in corpus: Limited to three disciplines (psychology, economics, political science) and studies with available data; Measurement bias: Conversion of diverse effect size estimates to standardized Cohen's d may not hold equally across analytical settings; Information context bias: Provision of varying amounts of methodological information (full text vs. abstract-only) could influence LLM behavior; Prompt engineering effects: Different prompt framings tested (neutral, confirmatory, critical) could introduce bias; Benchmark comparison bias: Human reanalysts pursued defensible alternative specifications, while LLM was instructed toward paper-anchored path

Limitations

  • The authors state: "First, the performance may depend on the modeling choices, such as the model and prompt
  • Our sensitivity analyses suggest that the main results are largely stable across these choices
  • Second, a key concern is training data contamination
  • Because the studies in our corpus were published before the training cutoff of the models, LLMs may have encountered some of the studies during pretraining, which could inflate reproducibility rates." They also note: "Third, our corpus is limited in size and to studies from psychology, economics, and political science
  • This reflects the difficulty and cost of assembling such datasets
  • even though our dataset is larger than others, this is a constraint that affects reproducibility audits broadly

Open questions raised

  • Whether LLMs can reliably carry out computational reproducibility assessments (addressed in this study)
  • Generalizability across different LLM models and architectures
  • Subtler effects of training data exposure beyond direct memorization
  • Expansion of corpus to broader disciplines beyond psychology, economics, and political science
  • Testing whether confirmed reproducibility extends to replication with newly collected data
  • Investigation of Goodhart-style incentives where authors optimize for automated screens
Data: SCORE project dataset (Systematizing Confidence in Open Research and Evidence): https://www.cos.io/score - 180 published empirical studies from psychology, economics, and political science with original data or replication data; Multi100 subset: 84 studies from large-scale human reanalysis effort; SCORE project (Systematizing Confidence in Open Research and Evidence): https://www.cos.io/score; Multi100 collaboration dataset (subset from reference [2]); Study data from original publications retrieved through SCORE repositories; SCORE (Systematizing Confidence in Open Research and Evidence) project dataset: https://www.cos.io/score; Multi100 dataset: Studies and reanalysis data from Ref. [2]; Evaluation corpus of 180 studies with predefined claims (publicly available via authors' code repository, specific URL not provided in text)Code: Code repository mentioned but URL not fully specified in paper: "(see our code repository)" - specific GitHub/repository URL not provided in accessible text; Code repository mentioned as available but specific URL not provided in paper: "see our code repository" (referenced in methods section); Authors' code repository mentioned but specific URL not provided in main text; Multi100 conversion code: https://github.com/marton-balazs-kovacs/multi100/blob/47c0b8c6dd68e19eb80fa8843dce18f0d3655ae1/analysis/multi100_raw_processed.qmd#L160-L165; Inspect AI framework: https://github.com/UKGovernmentBEIS/inspect_aiExtracted from: pdfAgreement 38%

Explore related topics

Related papers