AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang, Jin-Ge Yao, Zheng Liu et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark design with controlled evaluation over a curated corpus of over three million arXiv papers.
Sample
N = 1000, 6 groups
Primary method
Accuracy metric for Deep Research (exact-match indicator function), Intersection over Union (IoU) for Wide Research (set-based precision-recall metric). Comparative analysis across models using these metrics. Test-time scaling analysis (pass@k for Deep Research, best@k IoU for Wide Research). Error analysis and taxonomy development. No inferential statistics or hypothesis testing reported.
Main result
The study found that "AutoResearchBench presents a severe challenge. Top-performing systems achieve less than 10% in both metrics: Claude-Opus-4.6 peaks at 9.39% accuracy in Deep Research, and Gemini-3.1-Pro-Preview reaches only 9.31% IoU in Wide Research. The majority of evaluated models, including open-source systems, score below 5%." This contrasts sharply with general web-browsing benchmarks, indicating that "scientific literature discovery is not merely a harder instance of web browsing, but a distinct capability frontier for autonomous research agents."
Reports effect sizes.
Research paradigm
Empirical-Computational
Author conclusions
"AutoResearchBench establishes a rigorous diagnostic foundation for developing the next generation of reasoning-driven academic agents. Comprising 1,000 expert-curated queries grounded in a controlled, contamination-resistant corpus of over 3 million full-text papers, AutoResearchBench explicitly tests the intersection of long-horizon document browsing and complex scientific reasoning across Deep and Wide Research paradigms. Our comprehensive evaluation of frontier foundation models and end-to-end systems reveals a severe performance gap in real-world academic tasks: state-of-the-art performance peaks at a mere 9.39% accuracy for Deep Research and 9.31% IoU for Wide Research. By providing a clean evaluation ecosystem that exposes the insufficiency of shallow heuristic matching and general web navigation, AutoResearchBench establishes a rigorous diagnostic foundation for developing the next generation of reasoning-driven academic agents."
Risk of bias
Selection bias in target paper selection (preferentially focused on 'under-exposed yet high-quality works' with 10-100 citations); Annotation bias: Heavy reliance on human annotators for task construction and verification; Model selection bias: Evaluation models represent specific frontier systems that may share common limitations; Corpus bias: Fixed static corpus of arXiv papers may not reflect evolving scientific literature; Constraint construction bias: Full-text-first pipeline with deliberate obfuscation may not reflect natural research queries; Selection bias: Tasks constructed from curated corpus of over 3 million arXiv papers (primarily computer science domain); Annotation bias: Human experts involved in task construction and verification, potential for subjective constraint interpretation; Data contamination risk: Benchmark built on publicly accessible arXiv papers, potential overlap with training data of evaluated models; Domain bias: Limited to 8 core computer science domains, not representative of broader scientific literature; Task construction bias: Full-text-first pipeline may favor certain types of evidence over others; Potential incompleteness of ground-truth answer sets for Wide Research tasks may unfairly penalize models that retrieve valid but unannotated papers; potential data contamination from memorization effects mitigated through full-text-first task construction; annotation bias addressed through multi-stage verification pipeline and human expert auditing
Limitations
- The authors state that "AutoResearchBench is intentionally designed as a scientific papers discovery benchmark, but this scope also limits its coverage: it currently focuses on computer science papers in a fixed corpus and mainly evaluates text-based search and reasoning, leaving cross-domain science, multi-modal evidence, and continually evolving literature for future study
- In addition, although we adopt rigorous verification, exhaustive answer sets for wide search may still be challenging at the boundary of large corpora."
Open questions raised
- Scientific literature discovery requires distinct capabilities beyond general web browsing
- Current frontier agents lack weak scientific reasoning abilities
- Incomplete use of paper-level information in full-text analysis
- Difficulty handling long conjunctive queries with multiple constraints
- Insufficient comprehensiveness in set discovery for exhaustive literature coverage
- Limited reflection and iteration during search processes
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations