From Passive Generation to Investigation: A Proactive Scientific Peer Review Agent
Haishuo Fang, Yue Feng, Iryna Gurevych · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational experiment combining supervised fine-tuning and reinforcement learning (GRPO) on a dataset of 5,011 paper-review pairs from ICLR 2025/2026.
Sample
N = 5011, 6 groups
Primary method
Evaluation via three independent LLM judges with aggregated scores. Human evaluation using pairwise comparisons analyzed with Bradley-Terry model fitted over full matchup data. Inter-annotator agreement measured using Krippendorff's α, Fleiss' κ, and quadratic-weighted Cohen's κ2. Mean absolute error (MAE) used for score alignment computation. Avg-of-4 and best-of-4 scores reported from four independent review generations per method.
Main result
ProReviewer achieves the highest average score across five quality dimensions, with "ProReviewer with an 8B backbone, trained by supervised fine-tuning and optimized by reinforcement learning, achieves the highest average score across five quality dimensions, outperforming prompt-based methods with much larger frontier LLMs by up to 39% and the strongest fine-tuned baseline by 16% relatively." Specifically, ProReviewer (Qwen3-8B) achieves an overall score of 0.57 in average-of-4 evaluation and 0.65 in best-of-4, outperforming AI-Scientist-v2 (Qwen3.5-397B-A17B) which scores 0.45, and excels in Grounding (0.64) and Technical Depth (0.48).
Reports effect sizes and confidence intervals.
Research paradigm
Empirical computational research with reinforcement learning optimization
Author conclusions
"We introduced ProReviewer, a review agent that shifts automated peer review from passive generation to proactive investigation by formulating the review process as an MDP guided by a structured review log... Experiments show that ProReviewer with an 8B backbone outperforms both prompt-based systems with much larger frontier LLMs and fine-tuned baselines across automatic and human evaluation, while further analyses confirm its ability to detect cross-section inconsistencies and maintain robust performance on longer papers. These results suggest that proactive investigation supported by evidence tracking is a promising direction for LLM-assisted peer review and potentially for tasks requiring multi-step analytical reasoning over complex documents."
Risk of bias
Potential data contamination mitigated by temporal separation (test set from ICLR 2026, postdates model knowledge cutoff); LLM judge bias risk mitigated by using three diverse judges (GPT-5.4 nano, DeepSeek-V4 flash, RevUtil); Selection bias in test set limited to ICLR papers only (may not generalize to other domains); LLM-as-a-judge evaluation may exhibit systematic bias toward particular writing styles or model families; Single domain (ICLR conference papers) limits generalizability across scientific fields; Evaluation judges not used as base models in baselines, but potential model family preferences remain; Test set temporal separation (ICLR 2026 papers postdating base model cutoff) reduces but does not eliminate contamination risk; Domain specificity: trained and evaluated exclusively on ICLR papers, limiting generalizability to other fields; Potential judge bias: Despite using three independent LLM judges, systematic biases toward particular model families or writing styles possible; Selection bias in human evaluation: 50 papers randomly sampled from test set; evaluators may have implicit preferences; Data contamination mitigation: temporal separation used (test papers post-date base model knowledge cutoff), but some risk remains; Single-domain evaluation: ICLR papers may not represent all academic fields or review practices
Limitations
- The authors state that "the current implementation is text-only: the agent cannot directly inspect figures, which could include complementary evidence that is not accurately described in the text by the authors." Additionally, "ProReviewer is trained and evaluated on AI conference papers (ICLR), as other fields currently lack sufficient publicly available, clean manuscript–review pairs." Furthermore, "the current implementation focuses on intra-manuscript reasoning and does not perform external novelty search."
Open questions raised
- Extending the agent with multimodal perception to inspect figures and verify visual claims
- Adapting the approach to domains beyond AI conference papers (biomedicine, social sciences) once review data becomes available
- Adding external novelty search capabilities for cross-document reasoning
- The authors note the MDP formulation naturally accommodates these extensions by adding corresponding actions
- Multimodal perception to verify visual claims (figures, plots)
- Adaptation to domains beyond AI conference papers (biomedicine, social sciences) pending review data availability
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations