12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome

Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmarking study with multi-layered evaluation framework.

Sample

N = 13, 19 groups

Primary method

Pearson correlation coefficient between Process and combined outcome score; Kendall's τ (tau) for rank correlation between human judgments and MiroEval rankings; Fleiss' κ (kappa) for inter-rater agreement on validity and non-triviality (κ > 0.74); Linear regression analysis (Figure 5: synthesis quality vs. factuality, statement volume vs. precision); Standard deviation analysis for intra-judge stability testing; Majority-vote precision calculation for benchmark quality verification (92.0%)

Main result

The study found that "system rankings shift substantially across synthesis quality, factual precision, and research process rigor, demonstrating that each dimension provides non-redundant information." Additionally, "process quality serves as a reliable predictor of overall outcome while also revealing weaknesses invisible to output-level metrics, such as insufficient analytical depth and a significant traceability gap between reports and their underlying research procedures." The paper further demonstrates that "multimodal tasks pose substantially greater challenges, with most systems declining by 3 to 10 points," and that "the MiroThinker series demonstrates the most balanced performance, with MiroThinker-H1 achieving the highest overall scores in both text-only (77.5) and multimodal (74.5) settings."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/Positivist - measurement-based evaluation of computational systems

Author conclusions

"We introduced MiroEval, a benchmark and evaluation framework for deep research systems, comprising 100 tasks (70 text-only and 30 multimodal) assessed through three complementary layers: adaptive synthesis quality, agentic factuality, and process-centric evaluation. Our experiments across 13 leading systems show that the three dimensions capture complementary aspects of system capability; that process quality reliably predicts overall outcome while revealing weaknesses invisible to output-level metrics; and that multimodal tasks pose substantially greater challenges. Human verification confirms benchmark quality at 92.0% precision, and extensive robustness experiments together with a human ranking study (Kendall's τ = 0.91) validate the reliability of the evaluation framework. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents."

Risk of bias

Selection bias: benchmark drawn from internal testing patterns of single company (MiroMind), may not represent full diversity of real-world deep research queries; Evaluator bias: LLM-based evaluation using GPT-5.1, GPT-5.2, and GPT-5-mini as judges may introduce systematic judge model biases; Measurement bias: automated filtering pipeline may systematically exclude certain query types or domains; Conflict of interest: paper evaluates multiple MiroThinker variants (MiroThinker series) which show top performance, suggesting potential internal system favoritism; Temporal bias: benchmark snapshot taken March 2026; information landscape may have shifted for auto-generated queries; Multimodal evaluation bias: attachment processing relies on hybrid strategy (native + retrieval-augmented) which may favor certain system architectures; Judge model bias: LLM-based evaluation may encode model-specific biases; partial mitigation through cross-judge robustness checks (Gemini vs. GPT); Selection bias in query curation: User-derived queries sourced from internal testing phase only; generalizability to broader user populations unclear; Report generation timing: All reports collected in March 2026 within controlled window, reducing temporal confounders but limiting generalizability across seasons; Evaluator bias: Three human annotators conduct verification; while inter-annotator agreement reported (κ > 0.74), potential for systematic evaluator effects; Synthetic query generation: Trend-grounded queries anchored in documented web events may bias toward well-documented topics, potentially underrepresenting niche research needs; Judge model selection bias: Two different GPT versions used (GPT-5.1 for synthesis, GPT-5.2 for process, GPT-5-mini for factuality), though cross-judge consistency testing mitigates this; Evaluation prompt sensitivity: Though noted as minimal (<2 points), prompt design could influence results; Task construction bias: User-derived queries based on closed internal testing phase; real user patterns may differ; Multimodal evaluation limitations: Framework may not fully capture all failure modes in visual understanding; System access bias: Some systems lack multimodal support, creating incomplete comparisons

Open questions raised

  • Process evaluation currently limited to systems exposing intermediate reasoning traces; applicability to closed-source systems requiring improvement
  • Factuality evaluation identifies cross-source conflicts but does not resolve them; determining authoritative source when evidence disagrees remains unsolved
  • Current deep research systems exhibit universal weakness in Analytical Depth and Efficiency; both identified as primary intrinsic bottlenecks
  • Significant traceability gap exists: Report→Process scores substantially lower than Process→Report scores, indicating reports contain unsupported synthesis
  • Multimodal benchmarks remain limited; need for expanded multimodal evaluation beyond current 30-task multimodal subset
  • Temporal relevance: static benchmarks risk becoming stale; need for continuous refresh mechanisms as information landscape evolves
Data: Not explicitly listed as publicly available. Benchmark tasks and data appear to be proprietary or restricted: "all data handling follows strict confidentiality protocols: raw queries are processed only on access-controlled internal infrastructure." No URL or access information provided for the 100-task benchmark or evaluation results.Code: Not mentioned in documentExtracted from: pdfAgreement 60%

Explore related topics

Related papers