PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs
Yanjun Zhao, Tianxin Wei, Jiaru Zou, Xuying Ning, Yuanchen Bei, Lingjie Chen et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark construction and empirical evaluation.
Sample
> 1000, 11 groups
Primary method
F1 score calculation for task performance evaluation. LLM-as-a-Judge metric using GPT-4o with five-level rating scale and rule-based evaluation script. Accuracy calculated as proportion of predictions receiving score ≥4. Ablation studies examining maximum tool calls impact. Domain-wise and variance analysis of F1 scores and LLM-as-a-Judge ratings. Error taxonomy classification and failure mode analysis. Tool interaction depth and usage frequency analysis.
Main result
The benchmark reveals substantial performance disparities across models. "For Multimodal Ground, Gemini 2.5 Pro achieves the strongest overall performance, outperforming comparably sized Claude models by 20.4% in F1 score and 9.6% in LLM-as-a-Judge ratings." Additionally, "increasing the tool budget from 4 to 6 leads to the most substantial performance gains across models," and "explicitly providing source information leads to consistent performance improvements across evaluation metrics" with "Gemini-2.5-Pro model achieves improvements of 22.6% and 25.4% on the F1 score and the LLM-as-a-judge metric, respectively."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-Quantitative
Author conclusions
"We present the PaperMind benchmark for comprehensive scientific paper understanding that evaluates LLM-based systems across four interdependent task families: Multimodal Ground; Experimental Interpretation; Cross-Source Evidence Reasoning and Critical Assessment." The authors further state that "compared with existing benchmarks, our proposed benchmark moves beyond isolated retrieval and summarization to assess higher-level reasoning required for scientific research workflows" and conclude that "we believe this work facilitates more systematic evaluation and development of agentic systems for scientific literature understanding."
Risk of bias
Reliance on LLM-as-a-Judge for evaluation introduces potential bias in consistency and stability of judgments; Papers sourced from specific open-access repositories may not represent full distribution of scientific literature; Selection bias in peer-review filtering - only reviewer questions that elicit substantive author responses were retained; Model evaluation may reflect biases in pretraining data of evaluated LLMs; Reliance on LLM-as-a-Judge for evaluation may introduce bias and lacks human validation; Papers filtered by page length (5-20 pages) may exclude important short or long-form research; Papers with severe formatting issues excluded, potentially biasing toward well-formatted papers; Domain imbalance possible - papers collected from open-access sources which may not represent all scientific domains equally; Peer review data from OpenReview (computer science domain primarily) limits Critical Assessment task to CS domain; Question-answer pair construction uses Gemini 2.5 Pro as the generating model, potentially introducing model-specific biases; LLM-as-a-Judge evaluation bias: potential inconsistency and instability across diverse question types and domains; Domain representation bias: papers collected from open-access sources may not represent all scientific domains equally; Selection bias in paper filtering: papers removed for being too short (<5 pages) or too long (>20 pages) may exclude certain types of research; Bias propagation from underlying models: automated systems may propagate biases present in original academic literature or pretrained models; Tool preference bias: models demonstrated preference for general web search over specialized academic retrievers (arXiv_retriever)
Limitations
- The authors acknowledge that "although we rely exclusively on LLM-as-a-Judge for evaluation, the consistency and stability of such automated judgments remain an open challenge, especially across diverse question types." They further note that "the alignment between LLM-based judgments and human preferences may vary across different question types and domains, indicating room for improvement in evaluation stability and granularity."
Open questions raised
- Consistency and stability of LLM-as-a-Judge evaluation across diverse question types remains an open challenge
- Alignment between LLM-based judgments and human preferences varies across different question types and domains
- Limited understanding of agentic reasoning behaviors in realistic scientific workflows compared to static benchmarks
- Need for deeper analysis of tool-use failure patterns in scientific reasoning tasks
- The authors identify that existing benchmarks typically focus on isolated aspects of scientific QA (single-document comprehension, factual retrieval, or narrowly defined reasoning skills) rather than comprehensive integrated understanding. They note the need for evaluation of realistic scientific workflows involving tool use, evidence synthesis, and critical assessment. The paper also highlights that evaluation challenges remain: the consistency and stability of LLM-based judgments need improvement, and alignment between automated judgments and human preferences varies across question types and domains.
- Evaluation stability and granularity: alignment between LLM-based judgments and human preferences varies across different question types and domains
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations