ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
Bang H. Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage, Zack Ranjan, Sai Koneru et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark evaluation study with agent-based computational experiments.
Sample
N = 19, 4 groups
Primary method
Automated evaluation using LLMEval (GPT-4o as judge); leave-one-out cross-validation for extraction stage; macro and micro aggregation for data search metrics (precision, recall, F1, hit@any, hit@all); binary classification metrics (precision, recall, F1) for replication outcome; Spearman and Kendall correlation coefficients to validate LLMEval against human assessment; macro aggregation to treat binary outcome classes (Criteria Met/Unmet) equally.
Main result
The study found that "o3 and GPT-5 have the strongest computational performance for the execution stage" while "agents struggle in earlier stages that require locating replication data on the Internet." Additionally, "successful execution does not always translate into correct interpretation. Even when agents are able to reach beyond the generation stage and produce numerical results, interpretation errors, deviations from the pre-registered plan, and subtle implementation differences can lead to incorrect decisions."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical (computational benchmarking)
Author conclusions
"ReplicatorBench highlights promising directions for the development and evaluation of AI research assistants. Because current models struggle to locate appropriate data resources, we call for future work on agent development, specifically focusing on the planning stages, developing more specialized tools and effective search strategies to construct new replication samples." The authors conclude that "although state-of-the-art LLMs are often capable of performing complex computational workflows and iteratively resolving execution failures, this performance does not consistently yield correct replication judgments."
Risk of bias
Small sample size (n=19 studies) limits generalizability; LLMEval evaluation metric may introduce bias in scoring open-ended agent responses; Selection bias: filtering criteria eliminated many SCORE cases, keeping only 19 studies where focal claim testable with single quantitative check and clear pass/fail criteria; Limited to observational studies with online/archival data sources; experimental replication excluded; Potential language bias: LLMs may have preference for Python over other languages (R, etc.); Small sample size (n=19) may limit generalizability; LLMEval as automated judge may introduce systematic biases in grading; Selection bias in papers included (required data availability, single quantitative checks, observational studies); Potential LLM model-specific biases toward certain programming languages (Python preference noted); Potential bias in LLMEval scoring: the authors acknowledge treating LLMEval scores as approximations rather than absolute measures; Limited dataset: only 19 instances, constrained by scarcity of high-quality expert-documented replication efforts; Language preference bias: agents show differential performance across programming languages (Python vs. native), potentially affecting replication fidelity; Temporal bias: evaluation uses models with different release dates (GPT-4o May 2024, o3 Apr 2025, GPT-5 Aug 2025)
Limitations
- The authors state "the benchmark is constructed from a sample of 19 replication studies
- This scale is constrained by the scarcity of high-quality, expert-documented replication efforts that span multiple research stages." Additionally, they note "we recognize the shortcomings of using LLM-as-a-judge (LLMEval) for grading open-ended text, treating the rubric score as approximations rather than an absolute measure of replication competence." The benchmark also "focuses on observational studies in the SBS domains where data is web-retrievable, future work should develop benchmarks for experimental replication."
Open questions raised
- The authors identify several gaps: (1) lack of benchmarks assessing agents' capability to replicate when new data sample must be retrieved rather than readily available; (2) most existing benchmarks focus on final outcomes rather than the replication process itself; (3) limited evaluation of agents' behavior at each stage of replication; (4) need for benchmarks on experimental replication where agents must navigate controlled settings to collect or generate primary data; (5) need for better search strategies and tools to help agents locate replication data resources.
- Current LLM agents struggle with locating appropriate data resources for replication
- Lack of benchmarks for experimental replication (current work focuses on observational studies)
- Need for more specialized tools and search strategies for data sample construction
- Gap between execution success and correct interpretation/judgment in replication tasks
- Current models struggle to locate appropriate data resources; need for agent development focusing on planning stages and more specialized tools for data retrieval
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations