AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science
Abel Brodeur, David Valenta, Alexandru Marcoci, Juan P. Aparicio, Derek Mikola, Bruno Barbarioli et al. · Proceedings of the National Academy of Sciences · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1073/pnas.2524747123
Methodology & findings
Study design
Randomized controlled trial (RCT) with three treatment arms (human-only, AI-assisted, AI-led).
Sample
N = 288, 3 groups
Primary method
Ordinary least squares (OLS) regression as main analysis; logit and Poisson regressions as robustness checks; complementary Kaplan-Meier survival curves for time-to-event analysis (time to computational reproduction, time to first error detection); two-stage randomization procedure; exploratory subgroup analyses by AI experience and software type; thematic analysis of focus group data; pre-registered analyses with deviation tracking.
Main result
The study found that "Human-only and AI-assisted teams achieved comparable reproduction rates (94% vs. 91%) and performed similarly on most outcomes, except human-only teams identified significantly more major coding errors. Both substantially outperformed AI-led teams, which achieved only a 37% reproduction rate, detected fewer errors across all categories, proposed weaker robustness checks, and required more time." The research demonstrates that while AI assistance provides some efficiency benefits, human expertise remains critical for rigorous error detection and verification.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist
Author conclusions
"Human expertise remains critical to navigate challenges and provide interpretative guidance for reproducibility and error detection. The AI-assisted model—where humans work alongside AI tools—did not emerge as a winner over human-only teams in overall outcomes but outperformed AI-led teams on most of our outcomes. In scenarios where computational reproducibility, error detection, and robustness checks require in-depth understanding, domain knowledge, and flexible problem-solving, human involvement currently adds value. The ability to contextualize, interpret, and implement complex quantitative research remains a human strength, highlighting the limits of current AI in fully autonomous reproduction."
Risk of bias
Selection bias: Participants were coauthors on the paper, not a random sample of researchers; Hawthorne effect: Researchers aware of observation may have altered behavior; strong coding skills researchers may have exerted greater effort in human-only teams while responsibility may have been shifted to AI in AI-assisted/led settings; Attrition/task completion: Not all teams completed all tasks within 7 hours; Limited study sample: 12 studies used (with reuse across events), spanning limited range of social science methodologies; Demand characteristics and confirmation bias in focus groups: Participants were aware of headline results before focus group interviews; Selection bias: Participants were self-selected volunteers (coauthors) with doctoral degrees, not representative of all researchers; Hawthorne effect: Observation and professional identity may have influenced behavior, with researchers exerting greater effort in human-only teams; Attribution bias: Responsibility may have been shifted to AI in AI-assisted or AI-led settings; Model evolution bias: Multiple ChatGPT versions available during study period (GPT-4 through GPT-4o), potentially confounding temporal effects; Task complexity sampling bias: Limited to 12 studies with nonrandom selection; sample composition constrains generalizability; Integrity assumption: Relied on self-reported compliance of AI-led teams not to examine articles/code directly; Incentive bias: Coauthorship offered independent of performance, potentially reducing effort or introducing selection on motivation; Selection bias: Small, nonrandom set of 12 studies used; limited range of social science methodologies and replication difficulty levels; Hawthorne effect: Participant behavior influenced by observation and professional identity; researchers may have exerted greater effort in human-only teams or shifted responsibility to AI in other conditions; Incentive bias: Coauthorship offered independent of performance; no monetary compensation may have led to reduced effort for some teams, though this also reduced strategic behavior incentives; Attrition: Teams could leave before end of event if they believed tasks were completed; Confounding: Team composition, software preferences (Stata vs. R), and skill levels not fully controlled despite balancing across observables; AI experience heterogeneity: Participants had varying levels of AI familiarity, potentially affecting prompting quality and tool utilization; Confirmation bias in focus groups: Participants were aware of headline quantitative results when conducting focus groups conducted in April 2025
Limitations
- "One limitation is our sole focus on OpenAI's ChatGPT, meaning that we cannot generalize to all current AI models
- Furthermore, the limited timeframe of seven hours for study teams to complete their reproductions may not adequately reflect the conditions under which reproducibility efforts are conducted depending on the field of science
- In addition, participant incentives and attribution dynamics may have encouraged some teams to minimize time or effort, potentially increasing overreliance on AI tools
- Finally, our analysis is based on a small, nonrandom set of studies spanning a limited range of social science methodologies and replication difficulty levels." Additionally, "this sample composition constrains the extent to which our findings on AI assistance generalize across papers of different difficulty and across other scientific fields."
Open questions raised
- Future advancements in models optimized through reinforcement learning and chain-of-thought reasoning could improve reproduction of complex quantitative research
- Need for training models specifically in social science and quantitative research contexts rather than general-purpose models
- AI systems tailored for social science reproduction (e.g., with native support for R and Stata) could improve reproducibility outcomes
- Sustained benchmarking against humans needed as LLMs continue to evolve
- Research needed on potential for training AI systems with more sophisticated prompting strategies, parallel experimentation, and reduced labor costs through trained research assistants or undergraduate students
- Need for benchmarking of future AI-led reproducibility systems against human performance as models advance
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations