12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

Jiayu Wang, Weijiang Lv, Bowen Fu, Jing Fu, Jiayi Song, Lingyu Zhang et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical benchmark evaluation with systematic experimental design.

Sample

N = 82, 9 groups

Primary method

Descriptive statistics (success rates, pass rates, trajectory length distributions). Non-parametric analysis of task variance and failure patterns. Fine-grained unit test analysis with pass-rate deficits computed as differences between fine-grained and binary metrics.

Main result

The study found that "even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers." The results further indicate that "complex agent scaffolding is not a prerequisite for superior performance" and that minimalist agent architectures can outperform feature-heavy designs when paired with frontier models.

Reports effect sizes.

Research paradigm

Empiricist/Pragmatist - evaluates AI agent performance through benchmark experiments

Author conclusions

"In this work, we conceptualize the AARR benchmark series for evaluating LLM agents in authentic research scenarios. Specifically, we introduce AARRI-Bench, the inaugural benchmark in this series, and conduct extensive experiments across frontier models and agent harnesses. Our results show that despite recent advances in long-horizon agent capabilities, current systems still struggle with many subtle yet important details in real research workflows that remain straightforward for human researchers. We hope these findings can provide insights for future design, training and evaluation for agentic AI systems."

Risk of bias

Task design bias: Tasks drawn primarily from pain points of specific research team members; Selection bias: Frontier models and recent agentic systems may not represent full landscape; Evaluation bias: Heavy reliance on pattern matching in test code rather than LLM-as-judge; Potential publication bias: Only published/available models and harnesses evaluated; Task selection bias: Tasks were manually created by a single research team (diverse backgrounds acknowledged, but limited to one institution's researchers), potentially introducing systematic biases in what constitutes 'real research pain points'; Model availability bias: Evaluation limited to models accessible via official providers and OpenRouter; may not represent full landscape of available models; Harness selection bias: Only 3 agent harnesses tested; other architectures not evaluated; Evaluation metric bias: Binary reward metric may penalize partially correct solutions; fine-grained metrics employ pattern matching with acknowledged robustness limitations; Confounding: Model capability and harness design are interdependent factors; strong models may mask harness limitations or vice versa; Limited dataset scale may not represent full diversity of research tasks; Manual task construction by single research team may introduce selection bias toward team members' research domains; Exclusion of ultra-long-horizon tasks may bias evaluation toward simpler scenarios; Reliance on pattern matching in test code rather than LLM-as-judge may introduce evaluation brittleness

Limitations

  • "As the initial work in the AARR series, AARRI-Bench cannot achieve perfect balance across all aspects
  • Due to the limited human resources of our team, the dataset remains relatively small in scale
  • MCP and agent skills, supported though, have not yet been incorporated into the evaluation
  • Current tasks do not include ultra-long-horizon tasks, and almost all task evaluations are completed in less than ten minutes
  • To ensure high determinism and reproducibility of the evaluation, LLM-as-a-judge was not employed in AARRI, which required extensive pattern matching contents in the test code, compromising the robustness of the evaluation."

Open questions raised

  • Need for better understanding of research behavior and scientist-like qualities in AI agents
  • Limited evaluation of researcher qualities such as integrity, uncertainty awareness, and careful verification
  • Gap between tasks that are easy for humans but challenging for agents
  • Missing evaluation of ultra-long-horizon tasks
  • Need for integration of MCP and agent skills in future evaluations
  • Progression to AARRA and AARRS benchmarks for higher levels of research capability
Data: AARRI-Bench dataset; AARRI-Bench tasks released at https://github.com/AARR-bench/AARRI-bench; AARRI-BenchCode: AARRI-Bench; https://github.com/AARR-bench/AARRI-benchExtracted from: pdfAgreement 56%

Explore related topics

Related papers