EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
Xuan Dong, Huanyang Zheng, Tianhao Niu, Zhe Han, Pengzhan Li, Bofei Liu et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmarking study with multi-agent evaluation.
Sample
N = 102, 7 groups
Primary method
Agreement rate computation (simple proportion agreement), Cohen's κ for inter-rater reliability, error attribution using trace-based analysis with categorical classification, ablation studies with repeated measures across step budget variations. No hypothesis testing reported; evaluation is descriptive with benchmark-level metrics (ESR, Acc_pre, Acc_final, EC, MG).
Main result
The study found that "the best model reaching only 29.23% on the hard split, highlighting persistent bottlenecks in evidence grounding, multi-evidence fusion, and robust reuse of previously acquired evidence." Additionally, "even for the most capable closed-source systems, high turn-level accuracy on non-final turns does not translate into reliable end-to-end performance," with agents "often rely on incomplete visual evidence, link answers to incorrect figures or tables, and resort to erroneous memory or prior knowledge instead of seeking information that truly supports the evidence."
Reports effect sizes.
Research paradigm
Empirical-computational (benchmarking and agent evaluation)
Author conclusions
The authors conclude: "EpiBench will serve as a practical platform for developing more verifiable, efficient, and reproducible research agents." They emphasize that "developing more capable and reliable agents for this type of workflow is a highly valuable research direction, as even moderate improvements in task success could translate into substantial reductions in human labor and end-to-end completion time." The key conclusion is that "current systems remain far from reliable research assistance" and that "persistent bottlenecks in evidence grounding, multi-evidence fusion, and robust reuse of previously acquired evidence" must be addressed.
Risk of bias
Benchmark construction bias: Episodes were initially drafted by GPT-5.2, which could introduce model-specific linguistic or reasoning patterns, though substantial expert revision (92% of hard episodes underwent major restructuring) mitigates this concern; Evaluation judge bias: Primary evaluation used GPT-5.2 as judge, though reliability analysis with independent Gemini-2.5-Pro judge and human judges showed near-perfect agreement (98.7%-100.0%); Selection bias in seed papers: Benchmark constructed from 68 seed papers from public venues and preprint repositories, which may not represent the full diversity of scientific literature; Annotation bias: Five Ph.D. annotators curated episodes; disagreements resolved through discussion, but no inter-rater reliability statistics reported beyond consensus confirmation; Judge bias: While agreement analysis is reported (Table 4), the primary evaluation uses GPT-5.2 as the LLM judge, which may introduce bias toward GPT-5.2-generated outputs; Dataset construction bias: Episodes were initially generated by GPT-5.2, which could imprint specific patterns despite subsequent human curation; Selection bias: Benchmark composed of papers from 'six broad areas of computer vision and machine learning' and public venues/preprint repositories, potentially not representative of all research domains; Annotation bias: Curation performed by 'five annotators with Ph.D. degrees in computer science', potentially introducing domain-specific bias; Model evaluation bias: Tested models are primarily state-of-the-art commercial and recent open-source models; older or specialized models not included; Selection bias in benchmark construction: episodes drawn from classic papers in computer vision and machine learning may not represent broader scientific domains; Annotator bias: only five Ph.D.-level annotators involved in curation; potential for shared blind spots; Evaluation framework bias: benchmark design explicitly rewards certain behaviors (evidence reuse, avoiding prior knowledge) which may advantage certain model architectures; Judge bias: primary evaluation uses GPT-5.2 as judge; although inter-judge agreement is reported as high (98.7-100%), potential systematic bias remains; Small human baseline: only 2 experts limits human performance estimation; Model selection bias: choice of specific LMMs may not represent full landscape of capable systems
Open questions raised
- Process-level evaluation metrics for evidence reuse and evidence correctness in research agent benchmarks
- Benchmarks that comprehensively assess human-like workflows requiring proactive search, multi-turn interaction, and cross-paper multimodal evidence fusion
- Reliable multimodal evidence selection, indexing, and alignment in memory for research agents
- Better balancing of agent performance and interaction cost (token efficiency)
- Improved cross-paper evidence grounding and memory-based evidence reuse mechanisms
- Insufficient workflow-faithful benchmark design: existing benchmarks often not require proactive search or cross-paper multimodal integration
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations