Stop Automating Peer Review Without Rigorous Evaluation
Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed empirical methods including: (1) observational analysis of 75,800 real ICLR 2026 reviews with AI-generation labels; (2) controlled simulation experiments using AI reviewer agents on 60 randomly sampled papers; (3) paper laundering attacks using zero-shot LLM rewrites across 24 conditions (4 prompts × 2 launderer models × 3 reviewer models); (4) text similarity analysis using embeddings; (5) statistical testing (Welch's t-tests, Wilcoxon signed-rank tests, Cohen's d effect sizes)..
Sample
N = 75860, 11 groups
Primary method
Welch's t-tests for comparing group means; Wilcoxon signed-rank tests for paper laundering score differences; Cosine similarity using text embeddings (OpenAI text-embedding-3-small); Cohen's d for effect size quantification; Pearson correlation for AI-human score correlation (r = 0.15) and AI-AI correlation (r = 0.49); AUC analysis for predictive validity of review scores; N-gram analysis for phrase reuse identification
Main result
The study found two critical failures of current AI reviewing systems. First, "AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity," with AI-generated reviews showing "significantly higher within-group similarity (mean = 0.486) compared to other reviews (mean = 0.467; t = 3218, p < 0.0001, Cohen's d = 0.29)." Second, "AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers," with "zero-shot LLM rewrites boost AI review scores (+0.45, p < 0.0001) through stylistic modifications without human oversight."
Reports effect sizes and confidence intervals.
Research paradigm
Critical empiricism with normative positioning
Author conclusions
"Addressing the peer review crisis requires a science of peer review automation—not general-purpose LLMs deployed without rigorous evaluation." The authors argue that "today's AI systems should not produce paper reviews" and that "current AI reviewers fail both conditions" (preservation of review diversity and resistance to gaming). They propose that "addressing the peer review crisis requires a science of peer review automation, including rigorous evaluation of specific tools for specific tasks, not wholesale deployment of general-purpose LLMs."
Risk of bias
Detection accuracy of AI-generated reviews relying on third-party classification labels (Emi 2025) with potential misclassification errors; Simulation homogeneity bias: use of single fixed prompt across models may understate diversity compared to real-world varied prompts; Self-preference bias in LLMs: GPT reviewers show larger score increases than Claude when GPT is the launderer (Panickssery et al., 2024); Sample selection bias: 60 randomly sampled papers may not represent full diversity of paper types and quality levels at ICLR; Embedding metric limitations: cosine similarity on text embeddings may not capture argumentative or evaluative diversity; AI-generation label classification errors from Emi (2025); Self-preference bias in LLMs (documented with GPT reviewers showing larger score increases than Claude); Experimental homogeneity in simulation (single fixed prompt may underestimate diversity); Sample representativeness: 60 papers may not represent full diversity of paper types; Detection limitations: embedding-based metrics may not capture argumentative diversity; AI-generation label classification errors from Emi (2025) - though validated via independent checks; Self-preference bias in LLMs (GPT reviewers showed larger score increases than Claude); Experimental homogeneity in simulation (single fixed prompt, two models only); Embedding-based metrics may not capture true argumentative diversity; Sample limited to ICLR venue and 60 papers - generalizability concerns; Selection bias in paper sampling - randomly selected but may not represent full distribution
Limitations
- "Our AI reviewer simulations use only two models (GPT-5.1 and Claude Sonnet 4.5) with a single fixed prompt
- In practice, researchers and conferences may use diverse prompts, temperatures, and model versions, which could yield more varied outputs
- The high IntraSim thus partly reflects this experimental homogeneity rather than an intrinsic property of all possible AI reviewing setups." Additionally, "Our embedding-based similarity metrics capture linguistic and semantic patterns but do not directly measure diversity of viewpoints, arguments, or evaluative stances
- Two reviews could be linguistically similar yet offer different critiques, or vice versa." The "simulation experiments use 60 randomly sampled ICLR papers
- While sufficient to clearly show the two presented issues of gameability and non-diversity, this sample may not capture the full diversity of paper types and quality levels
- Additionally, our findings are specific to one venue (ICLR) and may not generalize to conferences with different review norms or paper distributions."
Open questions raised
- Need for adversarial robustness testing of AI systems in peer review before deployment
- Need for empirical studies measuring diversity of reviewer opinions when AI assistance is provided
- Need for user studies investigating how AI assistance affects reviewer behavior and whether humans catch AI errors
- Need for large-scale surveys understanding what different stakeholders value about peer review
- Need for research on how to enforce AI-usage policies in peer review
- Need for metrics that directly assess argumentative diversity rather than linguistic similarity
Explore related topics
Related papers
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Using AI to write scholarly publicationsMohammad Hosseini · 2023 · 264 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations