Generative Reviewer Agents: Scalable Simulacra of Peer Review
Nicolas Bougie, Narimawa Watanabe · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.emnlp-industry.8
Methodology & findings
Study design
Multi-round empirical evaluation framework.
Sample
N = 3797, 18 groups
Primary method
Bradley-Terry model for ranking preferences; quadratic-weighted Cohen's κ for ordinal score alignment; Pearson correlation coefficients for criterion alignment; t-tests for model performance comparison; binary cross-entropy loss minimization for BT coefficient estimation; win matrix construction; logistic regression. Results averaged over 20 runs (reviews) or 3 seeds (other experiments). Standard deviations reported as measures of variability.
Main result
The study found that GAR achieves an F1 score of 0.66 in predicting paper acceptance, "significantly exceeding human reviewers' 0.49 score" (p < 0.002). Furthermore, "GAR leads with a score of 0.684, outperforming human reviewers (0.523)" in pairwise preference comparisons using the Bradley-Terry model. The framework demonstrates that "GAR generated reviews stem from their depth, resulting in high preferences" through retrieval of relevant community-level reviews.
Reports effect sizes and confidence intervals.
Research paradigm
computational empiricism with agent-based simulation
Author conclusions
The authors conclude: "This research marks a step towards improving scientific writing and research by offering cost-effective, in-depth, and on-demand reviews." They further state: "Our vision is not to replace human reviewers but to enhance the review process by supporting them with synthetic reviewers capable of managing the increasing volume of submissions and providing early, constructive feedback. This collaboration between agents and human experts has the potential to accelerate scientific progress, foster innovation, and reduce time-to-publication."
Risk of bias
LLM training data bias: GAR may inherit biases from underlying LLMs (GPT-4o, Llama, Mistral); Institutional bias: potential preference for well-established institutions reflected in training data; Novelty detection bias: system overestimates low-novelty papers and underestimates highly novel papers; Domain-specific bias: evaluation limited to machine learning conferences; generalizability unknown; Information leakage: evaluation papers may have been partially included in LLM pre-training corpora; Memory contamination: title-matching filter to avoid review reuse is 'not foolproof' for modified manuscripts; Training data bias in LLMs affecting reviewer behavior simulation; Potential information leakage from pre-training corpora (papers may be in LLM training data); Institutional bias not fully mitigated (paper affiliation effects); Limited expertise representation in novelty assessment; Incomplete title/content matching for bias detection (fuzzy matching not implemented); Overestimation of low-novelty papers and underestimation of high-novelty papers; Lack of diversity in conference domains (ML conferences only); LLM training data bias - GAR may inherit biases from LLM training corpora; Information leakage - papers may have been included in LLM pre-training datasets; Institutional bias - though analysis suggests GAR may reduce rather than amplify institutional bias; Paradigm-shifting work detection - systematic bias against novel/unconventional ideas; Domain bias - trained/evaluated exclusively on machine learning conferences; Memory module contamination - risk of exact or near-duplicate papers biasing retrieval
Limitations
- The authors state: "First, our study primarily focuses on isolating and evaluating specific factors in the peer review process, such as reviewer dedication or expertise, instead of accounting for the inherent variability and arbitrariness that occur in real peer review scenarios." Additionally, "A limitation is GAR's difficulty in evaluating highly novel or paradigm-shifting work, as noted in Appendix A.4
- The system may struggle to recognize contributions that deviate from established norms, potentially overlooking groundbreaking ideas." Furthermore, "our evaluation was conducted in the context of machine learning conferences
- as such, the generalizability of our findings to other domains or conference communities remains to be established." The authors also note potential information leakage: "Since most frontier LLMs are trained on proprietary datasets, it remains difficult to ascertain whether evaluation papers may have been partially included in the models' training data."
Open questions raised
- Difficulty evaluating highly novel or paradigm-shifting work; future work should focus on equipping synthetic reviewers with ability to recognize novelty at a more nuanced level, including leveraging knowledge graph structure at community level or using citation embeddings
- Generalizability beyond machine learning conferences to other disciplines (mathematics, experimental sciences) with different review norms
- Handling sensitive information and intellectual property safeguards for practical deployment
- Advanced contamination detection strategies (fuzzy matching or content-based similarity filtering) to prevent reusing prior reviews or content
- Locally trained models on carefully controlled corpora to address information leakage from pre-training
- Difficulty recognizing genuinely groundbreaking or paradigm-shifting work
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations