12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

Zhihan Guo, Feiyang Xu, Yifan Li, Muzhi Li, Shuai Zou, Jiele Wu et al. · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
C
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2512.00986

Methodology & findings

Study design

Benchmark design and evaluation study.

Main result

The study reveals that "end-to-end academic DR remains a significant challenge for state-of-the-art agents" with "agents excel in different core modules, yet overall performance leaves significant room for improvement." Results show that "while Gemini stands out as the strongest overall due to its superior reasoning, its explicit planning capabilities are still nascent, and all agents, including the top retriever Grok, struggle significantly with retrieval on complex review papers." Additionally, "improving high-level planning capability is the crucial factor for unlocking the reasoning potential of foundational LLMs as backbones."

Research paradigm

Empirical-computational evaluation paradigm combining benchmarking with modular agent assessment

Author conclusions

"To address critical gaps in the evaluation of DR agents, we introduce ADRA-Bank, a human-annotated academic benchmark, and ADRA-Eval, a novel paradigm for the modular assessment of planning, retrieval, and reasoning. Our evaluation of state-of-the-art systems reveals a fragmented performance landscape where agents exhibit specialized strengths but also share critical weaknesses. By exposing these actionable failure modes, ADRA-Bank provides a diagnostic tool to guide the development of more reliable automatic academic research assistants."

Risk of bias

Paper selection bias toward high-citation (≥10 citations) and recent papers (post-2024), potentially underrepresenting emerging or specialized research areas; Disciplinary bias toward data-rich fields (Medicine, Biology) as evidenced by uneven performance across domains; Training data representation bias affecting agent performance across domains; Single annotator initial bias partially mitigated by dual-annotation protocol; Potential data leakage from model knowledge cutoffs analyzed but found to have negligible impact; Selection bias in paper curation: papers require ≥10 citations and publication after 2024, potentially underrepresenting emerging research; Domain representation bias: agents show uneven performance across domains (excel in Medicine/Biology, struggle in Finance/Materials), suggesting training corpus biases; Temporal leakage risks addressed through forward-citation filtering and knowledge cut-off validation; Human annotation bias mitigated through dual annotation with consensus adjudication (96.05% agreement, Cohen's Kappa 0.89); Selection bias in paper sources: papers required ≥10 citations and post-2024 publication, potentially excluding emerging or specialized topics; Disciplinary bias: uneven representation of academic domains in training corpora of evaluated models; Data leakage: models may have been trained on papers in benchmark corpus; authors analyzed temporal filtering but found minimal impact; Annotation quality variance: despite dual annotation and adjudication, consensus process may introduce systematic biases; Domain expert bias: annotation performed by senior PhD candidates who may have disciplinary biases

Limitations

  • "ADRA-Bank is designed as a diagnostic benchmark for academic deep research agents, with a focus on modularly evaluating planning, retrieval, and reasoning capabilities
  • While our dataset covers diverse academic disciplines and includes both research and review-oriented tasks, it is still a sampled benchmark rather than an exhaustive representation of all academic research scenarios
  • In particular, our paper selection criteria prioritize recent, publicly available, and impactful academic papers to ensure annotation quality, reproducibility, and verifiable provenance
  • As a result, highly specialized, emerging, or low-citation topics may be less represented."

Open questions raised

  • High-level planning capability as a critical bottleneck for foundational LLMs
  • Multi-source retrieval and breadth assessment on complex review papers
  • Cross-field consistency and domain-specific performance disparities
  • Need for targeted research to bridge capability gaps in underperforming disciplines (Finance, Materials, Earth Science)
  • Future extensions to broaden benchmark coverage of additional disciplines, publication types, and underrepresented research areas
  • Existing benchmarks focus narrowly on retrieval while neglecting high-level planning and reasoning
Data: ADRA-Bank: 200 instances across 10 academic domains (described in paper but availability link not provided in excerpt); Data sourced from: Web of Science database, CrossRef API, Semantic Scholar API; ADRA-Bank: 200 human-annotated instances across 10 academic domains (availability not explicitly stated in excerpt; DOI/GitHub repository not provided in accessible text); Source papers curated from Web of Science database with identifiers (DOIs, arXiv IDs, book titles, webpage links) compiled via CrossRef and Semantic Scholar APIs; ADRA-Bank: Human-annotated dataset of 200 instances across 10 academic domains (materials, finance, chemistry, computer science, medicine, biology, environmental science, energy, building and construction, earth science). Source papers retrieved from Web of Science and manually curated. Not explicitly stated if dataset will be released publicly.Code: Not explicitly stated in the paperExtracted from: pdfAgreement 43%

Explore related topics

Related papers