ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
Zhihan Guo, Feiyang Xu, Yifan Li, Muzhi Li, Shuai Zou, Jiele Wu et al. · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2512.00986
Methodology & findings
Study design
Benchmark design and evaluation study.
Main result
The study reveals that "end-to-end academic DR remains a significant challenge for state-of-the-art agents" with "agents excel in different core modules, yet overall performance leaves significant room for improvement." Results show that "while Gemini stands out as the strongest overall due to its superior reasoning, its explicit planning capabilities are still nascent, and all agents, including the top retriever Grok, struggle significantly with retrieval on complex review papers." Additionally, "improving high-level planning capability is the crucial factor for unlocking the reasoning potential of foundational LLMs as backbones."
Research paradigm
Empirical-computational evaluation paradigm combining benchmarking with modular agent assessment
Author conclusions
"To address critical gaps in the evaluation of DR agents, we introduce ADRA-Bank, a human-annotated academic benchmark, and ADRA-Eval, a novel paradigm for the modular assessment of planning, retrieval, and reasoning. Our evaluation of state-of-the-art systems reveals a fragmented performance landscape where agents exhibit specialized strengths but also share critical weaknesses. By exposing these actionable failure modes, ADRA-Bank provides a diagnostic tool to guide the development of more reliable automatic academic research assistants."
Risk of bias
Paper selection bias toward high-citation (≥10 citations) and recent papers (post-2024), potentially underrepresenting emerging or specialized research areas; Disciplinary bias toward data-rich fields (Medicine, Biology) as evidenced by uneven performance across domains; Training data representation bias affecting agent performance across domains; Single annotator initial bias partially mitigated by dual-annotation protocol; Potential data leakage from model knowledge cutoffs analyzed but found to have negligible impact; Selection bias in paper curation: papers require ≥10 citations and publication after 2024, potentially underrepresenting emerging research; Domain representation bias: agents show uneven performance across domains (excel in Medicine/Biology, struggle in Finance/Materials), suggesting training corpus biases; Temporal leakage risks addressed through forward-citation filtering and knowledge cut-off validation; Human annotation bias mitigated through dual annotation with consensus adjudication (96.05% agreement, Cohen's Kappa 0.89); Selection bias in paper sources: papers required ≥10 citations and post-2024 publication, potentially excluding emerging or specialized topics; Disciplinary bias: uneven representation of academic domains in training corpora of evaluated models; Data leakage: models may have been trained on papers in benchmark corpus; authors analyzed temporal filtering but found minimal impact; Annotation quality variance: despite dual annotation and adjudication, consensus process may introduce systematic biases; Domain expert bias: annotation performed by senior PhD candidates who may have disciplinary biases
Limitations
- "ADRA-Bank is designed as a diagnostic benchmark for academic deep research agents, with a focus on modularly evaluating planning, retrieval, and reasoning capabilities
- While our dataset covers diverse academic disciplines and includes both research and review-oriented tasks, it is still a sampled benchmark rather than an exhaustive representation of all academic research scenarios
- In particular, our paper selection criteria prioritize recent, publicly available, and impactful academic papers to ensure annotation quality, reproducibility, and verifiable provenance
- As a result, highly specialized, emerging, or low-citation topics may be less represented."
Open questions raised
- High-level planning capability as a critical bottleneck for foundational LLMs
- Multi-source retrieval and breadth assessment on complex review papers
- Cross-field consistency and domain-specific performance disparities
- Need for targeted research to bridge capability gaps in underperforming disciplines (Finance, Materials, Earth Science)
- Future extensions to broaden benchmark coverage of additional disciplines, publication types, and underrepresented research areas
- Existing benchmarks focus narrowly on retrieval while neglecting high-level planning and reasoning
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations