AI Coding Agents Can Reproduce Social Science Findings
Meysam Alizadeh, Mohsen Mosleh, Fabrizio Gilardi, Atoosa Kasirzadeh, Joshua Tucker · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark evaluation study.
Sample
N = 54, 10 groups
Primary method
Descriptive statistics (proportions, percentages, accuracy metrics). Accuracy calculated as proportion of tasks/papers answered correctly. Performance assessed across three independent runs with mean and range reported. Stratified analyses by programming language (Python, R, Stata), training-data cutoff (pre- and post-cutoff dates), and data repository (OSF, GitHub, CodeOcean). No inferential statistical tests reported (no p-values, t-tests, etc.). Failure rates calculated as proportion of cases where code fails to complete or produce expected output.
Main result
The study found that "Claude Code substantially outperformed Codex at both the task and paper levels. At the task level, Claude Code achieved a mean accuracy of 93.4%, compared with 62.1% for Codex-a difference of 31.3 percentage points. This gap widened at the paper level, where a paper was considered fully reproduced only if all of its constituent tasks were answered correctly: Claude Code attained 78.0% paper-level accuracy versus 35.8% for Codex, a difference of 42.2 percentage points." The study also found that "Claude Code autonomously resolved such problems in every case, constructing revised, executable replication pipelines without human intervention; by contrast, Codex failed to produce an answer for 17.8% of tasks and 27.0% of papers."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist - computational reproducibility evaluation
Author conclusions
The authors conclude: "Taken together, our results suggest that frontier AI coding agents are beginning to function as reliable executors of established computational workflows in social science. Although current systems remain sensitive to task framing and contextual cues, their ability to autonomously interpret and execute complex replication pipelines marks a meaningful step toward automated support for scientific reproducibility. Careful benchmarking and methodological transparency will be essential as such systems become more integrated into scientific practice."
Risk of bias
Selection bias: Benchmark constructed only from papers with complete and reproducible replication materials, overestimating real-world reproducibility rates; Survivor bias: Only papers with available data and code availability statements included; Publication venue bias: Restriction to high-tier journals (Nature, Science, PNAS, leading disciplinary journals) introduces bias toward better-documented research; Language bias: Limited to R, Python, and Stata; excludes other programming languages; Data contamination risk: Though authors assess this and find low metadata recovery rates, cannot fully rule out partial exposure to benchmark papers during model training; Prompt sensitivity: Confirmatory framing of tasks substantially affects agent behavior, particularly on non-reproducible tasks; Selection bias: Benchmark includes only papers with available replication materials and successful manual reproduction; Survivorship bias: Only papers from leading journals with formal data/code requirements included; Potential data contamination: Papers published before agent training cutoff may be in training data (though metadata recovery analysis argues against this); Anonymization bias: Removal of metadata may alter reproducibility difficulty in unrealistic ways; Selection bias: benchmark focuses only on reproducible papers, likely overestimating real-world performance; Limited scope: covers only subset of social science methods (political science, psychology, sociology, communication); Task format bias: structured tasks (coefficient extraction, plot interpretation) may not capture full diversity of empirical workflows; Paper selection bias: restricted to papers from leading journals (Nature, Science, PNAS, disciplinary leaders) with explicit data/code availability statements; Potential data contamination: although addressed through anonymization and metadata inference testing; Repository bias: papers sourced from OSF, GitHub, and Dataverse—may not be representative of broader social science
Limitations
- The authors state: "First, because the benchmark focuses on results that reproduce with available materials, it likely overestimates performance relative to real-world research environments where replication packages are incomplete or poorly documented
- Second, although the benchmark spans multiple disciplines and programming languages, it covers only a subset of social science methods
- Third, the evaluation relies on structured task formats-such as coefficient extraction or plot interpretation-that capture key elements of reproduction but cannot represent the full diversity of empirical workflows."
Open questions raised
- Authors identify several future research directions: (1) Benchmarks incorporating partially reproducible or incomplete replication materials would better approximate real-world research environments; (2) Expanding evaluation to replication and robustness tasks, such as testing alternative model specifications or applying established methods to new datasets, would assess agents' capacity to support broader stages of the scientific workflow; (3) Evaluating whether AI coding agents can select appropriate methods and arrive at correct conclusions would be a natural extension, particularly given evidence that human researchers themselves struggle with this.
- Benchmarks incorporating partially reproducible or incomplete replication materials to better approximate real-world research environments
- Expanding evaluation to replication and robustness tasks such as testing alternative model specifications or applying established methods to new datasets
- Evaluating whether AI coding agents can select appropriate methods and arrive at correct conclusions
- Expansion to replication and robustness tasks—testing alternative model specifications or applying established methods to new datasets
- Evaluation of whether AI coding agents can select appropriate methods and arrive at correct conclusions
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations