Automated reproducibility assessments in the social and behavioral sciences using large language models
Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten, Felix Henninger, Stefan Rose, Sarah Ball et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational simulation study using an agentic large language model (LLM) pipeline.
Main result
The study found that "LLMs can reproduce a substantial share of published findings in the social and behavioral sciences" with "the LLM reached the same qualitative conclusion as the original study in 80% of cases, and recovered the original effect sizes (using a ±0.05 tolerance in Cohen's d) in 24% of studies." In the subset with human reanalyses, "the LLM reached the same qualitative conclusion as the original study in 95% of studies, similar to human reanalysts (83%), and the LLM recovered the original effect sizes using a ±0.05 tolerance in 40% of studies, again broadly similar to human reanalysts (28%)"
Research paradigm
Empirical/quantitative positivism with computational verification
Author conclusions
The authors conclude: "In sum, our findings suggest that LLMs can offer a path toward more scalable reproducibility assessment by automatically performing systematic reanalyses that are otherwise costly to conduct manually. Given the current capabilities and limitations of LLMs, such automated reanalysis should be viewed as a scalable screening tool rather than a substitute for expert judgment. Nevertheless, it could support the research community in quality control and provide a first-pass check for scientific journals, which may ultimately help improve rigor and reproducibility in empirical research." They further note: "Rather than replacing expert judgment, LLM-based reanalysis could support editors, reviewers, and meta-scientists by providing a scalable first-pass check of whether reported findings can be recovered from the available data and study materials."
Risk of bias
Training data contamination: LLMs may have encountered study materials during pretraining; Memorization risk: Model may implicitly leverage prior exposure to studies; Model bias toward original reported results: LLM estimates correlated more strongly with original findings (r=0.46) than with human reanalyses (r=0.11); Selection bias: Corpus limited to studies with available data and extractable claims; Measurement bias: Cohen's d conversions assume standardized effect size comparability across heterogeneous analytical settings; Training data contamination: LLMs may have encountered published studies during pretraining, potentially inflating reproducibility rates; Selection bias: Corpus limited to studies from psychology, economics, and political science with available data; Anchoring bias: LLM pipeline may be biased toward recovering originally reported results rather than exploring alternative analytical paths; Prompt sensitivity: Different prompt framings could influence model behavior (though sensitivity analysis showed only 4 percentage point differences); Methodological assumptions: Conversion of heterogeneous effect sizes to Cohen's d rests on assumptions not universally applicable; Training data contamination: LLMs may have been exposed to studies during pretraining, potentially inflating reproducibility rates; Selection bias in corpus: Limited to three disciplines (psychology, economics, political science) and studies with available data; Measurement bias: Conversion of diverse effect size estimates to standardized Cohen's d may not hold equally across analytical settings; Information context bias: Provision of varying amounts of methodological information (full text vs. abstract-only) could influence LLM behavior; Prompt engineering effects: Different prompt framings tested (neutral, confirmatory, critical) could introduce bias; Benchmark comparison bias: Human reanalysts pursued defensible alternative specifications, while LLM was instructed toward paper-anchored path
Limitations
- The authors state: "First, the performance may depend on the modeling choices, such as the model and prompt
- Our sensitivity analyses suggest that the main results are largely stable across these choices
- Second, a key concern is training data contamination
- Because the studies in our corpus were published before the training cutoff of the models, LLMs may have encountered some of the studies during pretraining, which could inflate reproducibility rates." They also note: "Third, our corpus is limited in size and to studies from psychology, economics, and political science
- This reflects the difficulty and cost of assembling such datasets
- even though our dataset is larger than others, this is a constraint that affects reproducibility audits broadly
Open questions raised
- Whether LLMs can reliably carry out computational reproducibility assessments (addressed in this study)
- Generalizability across different LLM models and architectures
- Subtler effects of training data exposure beyond direct memorization
- Expansion of corpus to broader disciplines beyond psychology, economics, and political science
- Testing whether confirmed reproducibility extends to replication with newly collected data
- Investigation of Goodhart-style incentives where authors optimize for automated screens
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations