12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang, Jin-Ge Yao, Zheng Liu et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark design with controlled evaluation over a curated corpus of over three million arXiv papers.

Sample

N = 1000, 6 groups

Primary method

Accuracy metric for Deep Research (exact-match indicator function), Intersection over Union (IoU) for Wide Research (set-based precision-recall metric). Comparative analysis across models using these metrics. Test-time scaling analysis (pass@k for Deep Research, best@k IoU for Wide Research). Error analysis and taxonomy development. No inferential statistics or hypothesis testing reported.

Main result

The study found that "AutoResearchBench presents a severe challenge. Top-performing systems achieve less than 10% in both metrics: Claude-Opus-4.6 peaks at 9.39% accuracy in Deep Research, and Gemini-3.1-Pro-Preview reaches only 9.31% IoU in Wide Research. The majority of evaluated models, including open-source systems, score below 5%." This contrasts sharply with general web-browsing benchmarks, indicating that "scientific literature discovery is not merely a harder instance of web browsing, but a distinct capability frontier for autonomous research agents."

Reports effect sizes.

Research paradigm

Empirical-Computational

Author conclusions

"AutoResearchBench establishes a rigorous diagnostic foundation for developing the next generation of reasoning-driven academic agents. Comprising 1,000 expert-curated queries grounded in a controlled, contamination-resistant corpus of over 3 million full-text papers, AutoResearchBench explicitly tests the intersection of long-horizon document browsing and complex scientific reasoning across Deep and Wide Research paradigms. Our comprehensive evaluation of frontier foundation models and end-to-end systems reveals a severe performance gap in real-world academic tasks: state-of-the-art performance peaks at a mere 9.39% accuracy for Deep Research and 9.31% IoU for Wide Research. By providing a clean evaluation ecosystem that exposes the insufficiency of shallow heuristic matching and general web navigation, AutoResearchBench establishes a rigorous diagnostic foundation for developing the next generation of reasoning-driven academic agents."

Risk of bias

Selection bias in target paper selection (preferentially focused on 'under-exposed yet high-quality works' with 10-100 citations); Annotation bias: Heavy reliance on human annotators for task construction and verification; Model selection bias: Evaluation models represent specific frontier systems that may share common limitations; Corpus bias: Fixed static corpus of arXiv papers may not reflect evolving scientific literature; Constraint construction bias: Full-text-first pipeline with deliberate obfuscation may not reflect natural research queries; Selection bias: Tasks constructed from curated corpus of over 3 million arXiv papers (primarily computer science domain); Annotation bias: Human experts involved in task construction and verification, potential for subjective constraint interpretation; Data contamination risk: Benchmark built on publicly accessible arXiv papers, potential overlap with training data of evaluated models; Domain bias: Limited to 8 core computer science domains, not representative of broader scientific literature; Task construction bias: Full-text-first pipeline may favor certain types of evidence over others; Potential incompleteness of ground-truth answer sets for Wide Research tasks may unfairly penalize models that retrieve valid but unannotated papers; potential data contamination from memorization effects mitigated through full-text-first task construction; annotation bias addressed through multi-stage verification pipeline and human expert auditing

Limitations

  • The authors state that "AutoResearchBench is intentionally designed as a scientific papers discovery benchmark, but this scope also limits its coverage: it currently focuses on computer science papers in a fixed corpus and mainly evaluates text-based search and reasoning, leaving cross-domain science, multi-modal evidence, and continually evolving literature for future study
  • In addition, although we adopt rigorous verification, exhaustive answer sets for wide search may still be challenging at the boundary of large corpora."

Open questions raised

  • Scientific literature discovery requires distinct capabilities beyond general web browsing
  • Current frontier agents lack weak scientific reasoning abilities
  • Incomplete use of paper-level information in full-text analysis
  • Difficulty handling long conjunctive queries with multiple constraints
  • Insufficient comprehensiveness in set discovery for exhaustive literature coverage
  • Limited reflection and iteration during search processes
Data: AutoResearchBench benchmark dataset (1,000 queries with gold answers) - publicly released; Deep-Xiv corpus (3+ million arXiv papers with full-text extraction) - hosted on Deep-Xiv platform; AutoResearchBench benchmark: 1,000 curated queries (600 Deep Research, 400 Wide Research) - publicly released according to authors; DeepXiv corpus: Over 3 million arXiv papers with full-text extraction - hosted on Deep-Xiv platform; AutoResearchBench dataset with 1,000 curated queries (600 Deep Research, 400 Wide Research) over a corpus of 3+ million arXiv papers; authors state they "publicly release the dataset and evaluation pipeline"Code: Evaluation pipeline code - publicly released (repository URL not specified in text); Evaluation pipeline and benchmark infrastructure available for public release (specific GitHub/GitLab links not provided in text but promised in future release); Not explicitly stated in the paper; authors mention releasing evaluation pipeline publicly but specific repository URLs not providedExtracted from: pdfAgreement 50%

Explore related topics

Related papers