12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs

Yanjun Zhao, Tianxin Wei, Jiaru Zou, Xuying Ning, Yuanchen Bei, Lingjie Chen et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark construction and empirical evaluation.

Sample

> 1000, 11 groups

Primary method

F1 score calculation for task performance evaluation. LLM-as-a-Judge metric using GPT-4o with five-level rating scale and rule-based evaluation script. Accuracy calculated as proportion of predictions receiving score ≥4. Ablation studies examining maximum tool calls impact. Domain-wise and variance analysis of F1 scores and LLM-as-a-Judge ratings. Error taxonomy classification and failure mode analysis. Tool interaction depth and usage frequency analysis.

Main result

The benchmark reveals substantial performance disparities across models. "For Multimodal Ground, Gemini 2.5 Pro achieves the strongest overall performance, outperforming comparably sized Claude models by 20.4% in F1 score and 9.6% in LLM-as-a-Judge ratings." Additionally, "increasing the tool budget from 4 to 6 leads to the most substantial performance gains across models," and "explicitly providing source information leads to consistent performance improvements across evaluation metrics" with "Gemini-2.5-Pro model achieves improvements of 22.6% and 25.4% on the F1 score and the LLM-as-a-judge metric, respectively."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-Quantitative

Author conclusions

"We present the PaperMind benchmark for comprehensive scientific paper understanding that evaluates LLM-based systems across four interdependent task families: Multimodal Ground; Experimental Interpretation; Cross-Source Evidence Reasoning and Critical Assessment." The authors further state that "compared with existing benchmarks, our proposed benchmark moves beyond isolated retrieval and summarization to assess higher-level reasoning required for scientific research workflows" and conclude that "we believe this work facilitates more systematic evaluation and development of agentic systems for scientific literature understanding."

Risk of bias

Reliance on LLM-as-a-Judge for evaluation introduces potential bias in consistency and stability of judgments; Papers sourced from specific open-access repositories may not represent full distribution of scientific literature; Selection bias in peer-review filtering - only reviewer questions that elicit substantive author responses were retained; Model evaluation may reflect biases in pretraining data of evaluated LLMs; Reliance on LLM-as-a-Judge for evaluation may introduce bias and lacks human validation; Papers filtered by page length (5-20 pages) may exclude important short or long-form research; Papers with severe formatting issues excluded, potentially biasing toward well-formatted papers; Domain imbalance possible - papers collected from open-access sources which may not represent all scientific domains equally; Peer review data from OpenReview (computer science domain primarily) limits Critical Assessment task to CS domain; Question-answer pair construction uses Gemini 2.5 Pro as the generating model, potentially introducing model-specific biases; LLM-as-a-Judge evaluation bias: potential inconsistency and instability across diverse question types and domains; Domain representation bias: papers collected from open-access sources may not represent all scientific domains equally; Selection bias in paper filtering: papers removed for being too short (<5 pages) or too long (>20 pages) may exclude certain types of research; Bias propagation from underlying models: automated systems may propagate biases present in original academic literature or pretrained models; Tool preference bias: models demonstrated preference for general web search over specialized academic retrievers (arXiv_retriever)

Limitations

  • The authors acknowledge that "although we rely exclusively on LLM-as-a-Judge for evaluation, the consistency and stability of such automated judgments remain an open challenge, especially across diverse question types." They further note that "the alignment between LLM-based judgments and human preferences may vary across different question types and domains, indicating room for improvement in evaluation stability and granularity."

Open questions raised

  • Consistency and stability of LLM-as-a-Judge evaluation across diverse question types remains an open challenge
  • Alignment between LLM-based judgments and human preferences varies across different question types and domains
  • Limited understanding of agentic reasoning behaviors in realistic scientific workflows compared to static benchmarks
  • Need for deeper analysis of tool-use failure patterns in scientific reasoning tasks
  • The authors identify that existing benchmarks typically focus on isolated aspects of scientific QA (single-document comprehension, factual retrieval, or narrowly defined reasoning skills) rather than comprehensive integrated understanding. They note the need for evaluation of realistic scientific workflows involving tool use, evidence synthesis, and critical assessment. The paper also highlights that evaluation challenges remain: the consistency and stability of LLM-based judgments need improvement, and alignment between automated judgments and human preferences varies across question types and domains.
  • Evaluation stability and granularity: alignment between LLM-based judgments and human preferences varies across different question types and domains
Data: 3,000 scientific papers from ArXiv (https://arxiv.org); Papers from bioRxiv (https://www.biorxiv.org); Papers from Semantic Scholar (https://www.semanticscholar.org); Peer-review discussions from OpenReview (https://openreview.net/); PaperMind benchmark dataset (availability status not explicitly stated in text); 3,000 scientific papers from ArXiv, bioRxiv, and Semantic Scholar spanning agriculture, biology, chemistry, computer science, medicine, physics, and economics; Peer review discussions from OpenReview (294 reviewer-author QA pairs in computer science domain); Benchmark data sources: ArXiv (https://arxiv.org), bioRxiv (https://www.biorxiv.org), Semantic Scholar (https://www.semanticscholar.org), OpenReview (https://openreview.net/); PaperMind benchmark: 3,000 scientific papers from ArXiv, bioRxiv, and Semantic Scholar spanning seven domains (agriculture, biology, chemistry, computer science, medicine, physics, economics); Peer review discussions from OpenReview (294 reviewer-author QA pairs for Critical Assessment task)Code: smolagents framework (https://github.com/huggingface/smolagents) - used for tool interaction; smolagents framework (Roucher et al., 2025) - used for tool interaction and agent reasoning; paper2pdf tool (Clark and Divvala, 2016) - used for PDF to structured representation conversion; smolagents framework (Roucher et al., 2025) for tool-augmented reasoning; paper2pdf tool (Clark and Divvala, 2016) for PDF parsingExtracted from: pdfAgreement 55%

Explore related topics

Related papers