12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences

Bang H. Nguyen, Dominik Soós, Qian Ma, Rochana R. Obadage, Zack Ranjan, Sai Koneru et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark evaluation study with agent-based computational experiments.

Sample

N = 19, 4 groups

Primary method

Automated evaluation using LLMEval (GPT-4o as judge); leave-one-out cross-validation for extraction stage; macro and micro aggregation for data search metrics (precision, recall, F1, hit@any, hit@all); binary classification metrics (precision, recall, F1) for replication outcome; Spearman and Kendall correlation coefficients to validate LLMEval against human assessment; macro aggregation to treat binary outcome classes (Criteria Met/Unmet) equally.

Main result

The study found that "o3 and GPT-5 have the strongest computational performance for the execution stage" while "agents struggle in earlier stages that require locating replication data on the Internet." Additionally, "successful execution does not always translate into correct interpretation. Even when agents are able to reach beyond the generation stage and produce numerical results, interpretation errors, deviations from the pre-registered plan, and subtle implementation differences can lead to incorrect decisions."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical (computational benchmarking)

Author conclusions

"ReplicatorBench highlights promising directions for the development and evaluation of AI research assistants. Because current models struggle to locate appropriate data resources, we call for future work on agent development, specifically focusing on the planning stages, developing more specialized tools and effective search strategies to construct new replication samples." The authors conclude that "although state-of-the-art LLMs are often capable of performing complex computational workflows and iteratively resolving execution failures, this performance does not consistently yield correct replication judgments."

Risk of bias

Small sample size (n=19 studies) limits generalizability; LLMEval evaluation metric may introduce bias in scoring open-ended agent responses; Selection bias: filtering criteria eliminated many SCORE cases, keeping only 19 studies where focal claim testable with single quantitative check and clear pass/fail criteria; Limited to observational studies with online/archival data sources; experimental replication excluded; Potential language bias: LLMs may have preference for Python over other languages (R, etc.); Small sample size (n=19) may limit generalizability; LLMEval as automated judge may introduce systematic biases in grading; Selection bias in papers included (required data availability, single quantitative checks, observational studies); Potential LLM model-specific biases toward certain programming languages (Python preference noted); Potential bias in LLMEval scoring: the authors acknowledge treating LLMEval scores as approximations rather than absolute measures; Limited dataset: only 19 instances, constrained by scarcity of high-quality expert-documented replication efforts; Language preference bias: agents show differential performance across programming languages (Python vs. native), potentially affecting replication fidelity; Temporal bias: evaluation uses models with different release dates (GPT-4o May 2024, o3 Apr 2025, GPT-5 Aug 2025)

Limitations

  • The authors state "the benchmark is constructed from a sample of 19 replication studies
  • This scale is constrained by the scarcity of high-quality, expert-documented replication efforts that span multiple research stages." Additionally, they note "we recognize the shortcomings of using LLM-as-a-judge (LLMEval) for grading open-ended text, treating the rubric score as approximations rather than an absolute measure of replication competence." The benchmark also "focuses on observational studies in the SBS domains where data is web-retrievable, future work should develop benchmarks for experimental replication."

Open questions raised

  • The authors identify several gaps: (1) lack of benchmarks assessing agents' capability to replicate when new data sample must be retrieved rather than readily available; (2) most existing benchmarks focus on final outcomes rather than the replication process itself; (3) limited evaluation of agents' behavior at each stage of replication; (4) need for benchmarks on experimental replication where agents must navigate controlled settings to collect or generate primary data; (5) need for better search strategies and tools to help agents locate replication data resources.
  • Current LLM agents struggle with locating appropriate data resources for replication
  • Lack of benchmarks for experimental replication (current work focuses on observational studies)
  • Need for more specialized tools and search strategies for data sample construction
  • Gap between execution success and correct interpretation/judgment in replication tasks
  • Current models struggle to locate appropriate data resources; need for agent development focusing on planning stages and more specialized tools for data retrieval
Data: ReplicatorBench dataset: 19 instances from SCORE project (Systemizing Confidence in Open Research and Evidence), sourced from peer-reviewed journals in six subject categories in social and behavioral sciences. Available as part of the benchmark but specific URL not provided in paper.; ReplicatorBench dataset consisting of 19 replication instances with human expert reports from SCORE project; Replication data for each study (specific URLs/access methods not provided in abstract); ReplicatorBench: 19 replication instances from SCORE project papers; Replication data curated from online/archival sources (specific URLs documented in benchmark instances); Papers sampled from six subject categories in social and behavioral sciencesCode: Not explicitly mentioned with URL in the paper; ReplicatorAgent framework code (specific GitHub URL not provided in document); Code repository not explicitly provided in the paper; ReplicatorAgent framework described but no GitHub link providedExtracted from: pdfAgreement 41%

Explore related topics

Related papers