12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents

Xuan Dong, Huanyang Zheng, Tianhao Niu, Zhe Han, Pengzhan Li, Bofei Liu et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmarking study with multi-agent evaluation.

Sample

N = 102, 7 groups

Primary method

Agreement rate computation (simple proportion agreement), Cohen's κ for inter-rater reliability, error attribution using trace-based analysis with categorical classification, ablation studies with repeated measures across step budget variations. No hypothesis testing reported; evaluation is descriptive with benchmark-level metrics (ESR, Acc_pre, Acc_final, EC, MG).

Main result

The study found that "the best model reaching only 29.23% on the hard split, highlighting persistent bottlenecks in evidence grounding, multi-evidence fusion, and robust reuse of previously acquired evidence." Additionally, "even for the most capable closed-source systems, high turn-level accuracy on non-final turns does not translate into reliable end-to-end performance," with agents "often rely on incomplete visual evidence, link answers to incorrect figures or tables, and resort to erroneous memory or prior knowledge instead of seeking information that truly supports the evidence."

Reports effect sizes.

Research paradigm

Empirical-computational (benchmarking and agent evaluation)

Author conclusions

The authors conclude: "EpiBench will serve as a practical platform for developing more verifiable, efficient, and reproducible research agents." They emphasize that "developing more capable and reliable agents for this type of workflow is a highly valuable research direction, as even moderate improvements in task success could translate into substantial reductions in human labor and end-to-end completion time." The key conclusion is that "current systems remain far from reliable research assistance" and that "persistent bottlenecks in evidence grounding, multi-evidence fusion, and robust reuse of previously acquired evidence" must be addressed.

Risk of bias

Benchmark construction bias: Episodes were initially drafted by GPT-5.2, which could introduce model-specific linguistic or reasoning patterns, though substantial expert revision (92% of hard episodes underwent major restructuring) mitigates this concern; Evaluation judge bias: Primary evaluation used GPT-5.2 as judge, though reliability analysis with independent Gemini-2.5-Pro judge and human judges showed near-perfect agreement (98.7%-100.0%); Selection bias in seed papers: Benchmark constructed from 68 seed papers from public venues and preprint repositories, which may not represent the full diversity of scientific literature; Annotation bias: Five Ph.D. annotators curated episodes; disagreements resolved through discussion, but no inter-rater reliability statistics reported beyond consensus confirmation; Judge bias: While agreement analysis is reported (Table 4), the primary evaluation uses GPT-5.2 as the LLM judge, which may introduce bias toward GPT-5.2-generated outputs; Dataset construction bias: Episodes were initially generated by GPT-5.2, which could imprint specific patterns despite subsequent human curation; Selection bias: Benchmark composed of papers from 'six broad areas of computer vision and machine learning' and public venues/preprint repositories, potentially not representative of all research domains; Annotation bias: Curation performed by 'five annotators with Ph.D. degrees in computer science', potentially introducing domain-specific bias; Model evaluation bias: Tested models are primarily state-of-the-art commercial and recent open-source models; older or specialized models not included; Selection bias in benchmark construction: episodes drawn from classic papers in computer vision and machine learning may not represent broader scientific domains; Annotator bias: only five Ph.D.-level annotators involved in curation; potential for shared blind spots; Evaluation framework bias: benchmark design explicitly rewards certain behaviors (evidence reuse, avoiding prior knowledge) which may advantage certain model architectures; Judge bias: primary evaluation uses GPT-5.2 as judge; although inter-judge agreement is reported as high (98.7-100%), potential systematic bias remains; Small human baseline: only 2 experts limits human performance estimation; Model selection bias: choice of specific LMMs may not represent full landscape of capable systems

Open questions raised

  • Process-level evaluation metrics for evidence reuse and evidence correctness in research agent benchmarks
  • Benchmarks that comprehensively assess human-like workflows requiring proactive search, multi-turn interaction, and cross-paper multimodal evidence fusion
  • Reliable multimodal evidence selection, indexing, and alignment in memory for research agents
  • Better balancing of agent performance and interaction cost (token efficiency)
  • Improved cross-paper evidence grounding and memory-based evidence reuse mechanisms
  • Insufficient workflow-faithful benchmark design: existing benchmarks often not require proactive search or cross-paper multimodal integration
Data: EpiBench benchmark (102 episodic multi-turn research tasks) - availability not explicitly stated in paper; EpiBench benchmark: 102 episodes with expert annotations and evidence unit specifications (source corpus: 485 papers from OpenReview and arXiv); EpiBench: 102 episodic multi-turn tasks; availability status unclear from paper; Source corpus: 485 papers from OpenReview and arXivCode: smolagents [21] - adapted and used as base for agent implementation; smolagents (adapted from this open-source framework for the agent implementation); smolagents framework used and adapted (referenced as [21])Extracted from: pdfAgreement 51%

Explore related topics

Related papers