12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science

Abel Brodeur, David Valenta, Alexandru Marcoci, Juan P. Aparicio, Derek Mikola, Bruno Barbarioli et al. · Proceedings of the National Academy of Sciences · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1073/pnas.2524747123

Methodology & findings

Study design

Randomized controlled trial (RCT) with three treatment arms (human-only, AI-assisted, AI-led).

Sample

N = 288, 3 groups

Primary method

Ordinary least squares (OLS) regression as main analysis; logit and Poisson regressions as robustness checks; complementary Kaplan-Meier survival curves for time-to-event analysis (time to computational reproduction, time to first error detection); two-stage randomization procedure; exploratory subgroup analyses by AI experience and software type; thematic analysis of focus group data; pre-registered analyses with deviation tracking.

Main result

The study found that "Human-only and AI-assisted teams achieved comparable reproduction rates (94% vs. 91%) and performed similarly on most outcomes, except human-only teams identified significantly more major coding errors. Both substantially outperformed AI-led teams, which achieved only a 37% reproduction rate, detected fewer errors across all categories, proposed weaker robustness checks, and required more time." The research demonstrates that while AI assistance provides some efficiency benefits, human expertise remains critical for rigorous error detection and verification.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist

Author conclusions

"Human expertise remains critical to navigate challenges and provide interpretative guidance for reproducibility and error detection. The AI-assisted model—where humans work alongside AI tools—did not emerge as a winner over human-only teams in overall outcomes but outperformed AI-led teams on most of our outcomes. In scenarios where computational reproducibility, error detection, and robustness checks require in-depth understanding, domain knowledge, and flexible problem-solving, human involvement currently adds value. The ability to contextualize, interpret, and implement complex quantitative research remains a human strength, highlighting the limits of current AI in fully autonomous reproduction."

Risk of bias

Selection bias: Participants were coauthors on the paper, not a random sample of researchers; Hawthorne effect: Researchers aware of observation may have altered behavior; strong coding skills researchers may have exerted greater effort in human-only teams while responsibility may have been shifted to AI in AI-assisted/led settings; Attrition/task completion: Not all teams completed all tasks within 7 hours; Limited study sample: 12 studies used (with reuse across events), spanning limited range of social science methodologies; Demand characteristics and confirmation bias in focus groups: Participants were aware of headline results before focus group interviews; Selection bias: Participants were self-selected volunteers (coauthors) with doctoral degrees, not representative of all researchers; Hawthorne effect: Observation and professional identity may have influenced behavior, with researchers exerting greater effort in human-only teams; Attribution bias: Responsibility may have been shifted to AI in AI-assisted or AI-led settings; Model evolution bias: Multiple ChatGPT versions available during study period (GPT-4 through GPT-4o), potentially confounding temporal effects; Task complexity sampling bias: Limited to 12 studies with nonrandom selection; sample composition constrains generalizability; Integrity assumption: Relied on self-reported compliance of AI-led teams not to examine articles/code directly; Incentive bias: Coauthorship offered independent of performance, potentially reducing effort or introducing selection on motivation; Selection bias: Small, nonrandom set of 12 studies used; limited range of social science methodologies and replication difficulty levels; Hawthorne effect: Participant behavior influenced by observation and professional identity; researchers may have exerted greater effort in human-only teams or shifted responsibility to AI in other conditions; Incentive bias: Coauthorship offered independent of performance; no monetary compensation may have led to reduced effort for some teams, though this also reduced strategic behavior incentives; Attrition: Teams could leave before end of event if they believed tasks were completed; Confounding: Team composition, software preferences (Stata vs. R), and skill levels not fully controlled despite balancing across observables; AI experience heterogeneity: Participants had varying levels of AI familiarity, potentially affecting prompting quality and tool utilization; Confirmation bias in focus groups: Participants were aware of headline quantitative results when conducting focus groups conducted in April 2025

Limitations

  • "One limitation is our sole focus on OpenAI's ChatGPT, meaning that we cannot generalize to all current AI models
  • Furthermore, the limited timeframe of seven hours for study teams to complete their reproductions may not adequately reflect the conditions under which reproducibility efforts are conducted depending on the field of science
  • In addition, participant incentives and attribution dynamics may have encouraged some teams to minimize time or effort, potentially increasing overreliance on AI tools
  • Finally, our analysis is based on a small, nonrandom set of studies spanning a limited range of social science methodologies and replication difficulty levels." Additionally, "this sample composition constrains the extent to which our findings on AI assistance generalize across papers of different difficulty and across other scientific fields."

Open questions raised

  • Future advancements in models optimized through reinforcement learning and chain-of-thought reasoning could improve reproduction of complex quantitative research
  • Need for training models specifically in social science and quantitative research contexts rather than general-purpose models
  • AI systems tailored for social science reproduction (e.g., with native support for R and Stata) could improve reproducibility outcomes
  • Sustained benchmarking against humans needed as LLMs continue to evolve
  • Research needed on potential for training AI systems with more sophisticated prompting strategies, parallel experimentation, and reduced labor costs through trained research assistants or undergraduate students
  • Need for benchmarking of future AI-led reproducibility systems against human performance as models advance
Data: Not explicitly stated. The paper indicates that teams were provided with journal articles, online appendices as PDFs, original authors' replication packages, and screenshots of exhibits to reproduce. No public dataset repository is mentioned.Code: Not explicitly stated. AI conversation histories (ChatGPT transcripts) were required from AI-assisted and AI-led teams but storage location not specified.Extracted from: pdfAgreement 53%

Explore related topics

Related papers