MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmarking study with multi-layered evaluation framework.
Sample
N = 13, 19 groups
Primary method
Pearson correlation coefficient between Process and combined outcome score; Kendall's τ (tau) for rank correlation between human judgments and MiroEval rankings; Fleiss' κ (kappa) for inter-rater agreement on validity and non-triviality (κ > 0.74); Linear regression analysis (Figure 5: synthesis quality vs. factuality, statement volume vs. precision); Standard deviation analysis for intra-judge stability testing; Majority-vote precision calculation for benchmark quality verification (92.0%)
Main result
The study found that "system rankings shift substantially across synthesis quality, factual precision, and research process rigor, demonstrating that each dimension provides non-redundant information." Additionally, "process quality serves as a reliable predictor of overall outcome while also revealing weaknesses invisible to output-level metrics, such as insufficient analytical depth and a significant traceability gap between reports and their underlying research procedures." The paper further demonstrates that "multimodal tasks pose substantially greater challenges, with most systems declining by 3 to 10 points," and that "the MiroThinker series demonstrates the most balanced performance, with MiroThinker-H1 achieving the highest overall scores in both text-only (77.5) and multimodal (74.5) settings."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist - measurement-based evaluation of computational systems
Author conclusions
"We introduced MiroEval, a benchmark and evaluation framework for deep research systems, comprising 100 tasks (70 text-only and 30 multimodal) assessed through three complementary layers: adaptive synthesis quality, agentic factuality, and process-centric evaluation. Our experiments across 13 leading systems show that the three dimensions capture complementary aspects of system capability; that process quality reliably predicts overall outcome while revealing weaknesses invisible to output-level metrics; and that multimodal tasks pose substantially greater challenges. Human verification confirms benchmark quality at 92.0% precision, and extensive robustness experiments together with a human ranking study (Kendall's τ = 0.91) validate the reliability of the evaluation framework. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents."
Risk of bias
Selection bias: benchmark drawn from internal testing patterns of single company (MiroMind), may not represent full diversity of real-world deep research queries; Evaluator bias: LLM-based evaluation using GPT-5.1, GPT-5.2, and GPT-5-mini as judges may introduce systematic judge model biases; Measurement bias: automated filtering pipeline may systematically exclude certain query types or domains; Conflict of interest: paper evaluates multiple MiroThinker variants (MiroThinker series) which show top performance, suggesting potential internal system favoritism; Temporal bias: benchmark snapshot taken March 2026; information landscape may have shifted for auto-generated queries; Multimodal evaluation bias: attachment processing relies on hybrid strategy (native + retrieval-augmented) which may favor certain system architectures; Judge model bias: LLM-based evaluation may encode model-specific biases; partial mitigation through cross-judge robustness checks (Gemini vs. GPT); Selection bias in query curation: User-derived queries sourced from internal testing phase only; generalizability to broader user populations unclear; Report generation timing: All reports collected in March 2026 within controlled window, reducing temporal confounders but limiting generalizability across seasons; Evaluator bias: Three human annotators conduct verification; while inter-annotator agreement reported (κ > 0.74), potential for systematic evaluator effects; Synthetic query generation: Trend-grounded queries anchored in documented web events may bias toward well-documented topics, potentially underrepresenting niche research needs; Judge model selection bias: Two different GPT versions used (GPT-5.1 for synthesis, GPT-5.2 for process, GPT-5-mini for factuality), though cross-judge consistency testing mitigates this; Evaluation prompt sensitivity: Though noted as minimal (<2 points), prompt design could influence results; Task construction bias: User-derived queries based on closed internal testing phase; real user patterns may differ; Multimodal evaluation limitations: Framework may not fully capture all failure modes in visual understanding; System access bias: Some systems lack multimodal support, creating incomplete comparisons
Open questions raised
- Process evaluation currently limited to systems exposing intermediate reasoning traces; applicability to closed-source systems requiring improvement
- Factuality evaluation identifies cross-source conflicts but does not resolve them; determining authoritative source when evidence disagrees remains unsolved
- Current deep research systems exhibit universal weakness in Analytical Depth and Efficiency; both identified as primary intrinsic bottlenecks
- Significant traceability gap exists: Report→Process scores substantially lower than Process→Report scores, indicating reports contain unsupported synthesis
- Multimodal benchmarks remain limited; need for expanded multimodal evaluation beyond current 30-task multimodal subset
- Temporal relevance: static benchmarks risk becoming stale; need for continuous refresh mechanisms as information landscape evolves
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations