Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
Jiayu Wang, Weijiang Lv, Bowen Fu, Jing Fu, Jiayi Song, Lingyu Zhang et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical benchmark evaluation with systematic experimental design.
Sample
N = 82, 9 groups
Primary method
Descriptive statistics (success rates, pass rates, trajectory length distributions). Non-parametric analysis of task variance and failure patterns. Fine-grained unit test analysis with pass-rate deficits computed as differences between fine-grained and binary metrics.
Main result
The study found that "even the best-performing configuration (Mini-SWE-Agent with Claude Opus 4.7) achieves only 68.3% success rate, frequently overlooking subtle yet critical details that are obvious to real human researchers." The results further indicate that "complex agent scaffolding is not a prerequisite for superior performance" and that minimalist agent architectures can outperform feature-heavy designs when paired with frontier models.
Reports effect sizes.
Research paradigm
Empiricist/Pragmatist - evaluates AI agent performance through benchmark experiments
Author conclusions
"In this work, we conceptualize the AARR benchmark series for evaluating LLM agents in authentic research scenarios. Specifically, we introduce AARRI-Bench, the inaugural benchmark in this series, and conduct extensive experiments across frontier models and agent harnesses. Our results show that despite recent advances in long-horizon agent capabilities, current systems still struggle with many subtle yet important details in real research workflows that remain straightforward for human researchers. We hope these findings can provide insights for future design, training and evaluation for agentic AI systems."
Risk of bias
Task design bias: Tasks drawn primarily from pain points of specific research team members; Selection bias: Frontier models and recent agentic systems may not represent full landscape; Evaluation bias: Heavy reliance on pattern matching in test code rather than LLM-as-judge; Potential publication bias: Only published/available models and harnesses evaluated; Task selection bias: Tasks were manually created by a single research team (diverse backgrounds acknowledged, but limited to one institution's researchers), potentially introducing systematic biases in what constitutes 'real research pain points'; Model availability bias: Evaluation limited to models accessible via official providers and OpenRouter; may not represent full landscape of available models; Harness selection bias: Only 3 agent harnesses tested; other architectures not evaluated; Evaluation metric bias: Binary reward metric may penalize partially correct solutions; fine-grained metrics employ pattern matching with acknowledged robustness limitations; Confounding: Model capability and harness design are interdependent factors; strong models may mask harness limitations or vice versa; Limited dataset scale may not represent full diversity of research tasks; Manual task construction by single research team may introduce selection bias toward team members' research domains; Exclusion of ultra-long-horizon tasks may bias evaluation toward simpler scenarios; Reliance on pattern matching in test code rather than LLM-as-judge may introduce evaluation brittleness
Limitations
- "As the initial work in the AARR series, AARRI-Bench cannot achieve perfect balance across all aspects
- Due to the limited human resources of our team, the dataset remains relatively small in scale
- MCP and agent skills, supported though, have not yet been incorporated into the evaluation
- Current tasks do not include ultra-long-horizon tasks, and almost all task evaluations are completed in less than ten minutes
- To ensure high determinism and reproducibility of the evaluation, LLM-as-a-judge was not employed in AARRI, which required extensive pattern matching contents in the test code, compromising the robustness of the evaluation."
Open questions raised
- Need for better understanding of research behavior and scientist-like qualities in AI agents
- Limited evaluation of researcher qualities such as integrity, uncertainty awareness, and careful verification
- Gap between tasks that are easy for humans but challenging for agents
- Missing evaluation of ultra-long-horizon tasks
- Need for integration of MCP and agent skills in future evaluations
- Progression to AARRA and AARRS benchmarks for higher levels of research capability
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- What ChatGPT means for universities: Perceptions of scholars and studentsMehmet Fırat · 2023 · 405 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations