12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Stop Automating Peer Review Without Rigorous Evaluation

Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed empirical methods including: (1) observational analysis of 75,800 real ICLR 2026 reviews with AI-generation labels; (2) controlled simulation experiments using AI reviewer agents on 60 randomly sampled papers; (3) paper laundering attacks using zero-shot LLM rewrites across 24 conditions (4 prompts × 2 launderer models × 3 reviewer models); (4) text similarity analysis using embeddings; (5) statistical testing (Welch's t-tests, Wilcoxon signed-rank tests, Cohen's d effect sizes)..

Sample

N = 75860, 11 groups

Primary method

Welch's t-tests for comparing group means; Wilcoxon signed-rank tests for paper laundering score differences; Cosine similarity using text embeddings (OpenAI text-embedding-3-small); Cohen's d for effect size quantification; Pearson correlation for AI-human score correlation (r = 0.15) and AI-AI correlation (r = 0.49); AUC analysis for predictive validity of review scores; N-gram analysis for phrase reuse identification

Main result

The study found two critical failures of current AI reviewing systems. First, "AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity," with AI-generated reviews showing "significantly higher within-group similarity (mean = 0.486) compared to other reviews (mean = 0.467; t = 3218, p < 0.0001, Cohen's d = 0.29)." Second, "AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers," with "zero-shot LLM rewrites boost AI review scores (+0.45, p < 0.0001) through stylistic modifications without human oversight."

Reports effect sizes and confidence intervals.

Research paradigm

Critical empiricism with normative positioning

Author conclusions

"Addressing the peer review crisis requires a science of peer review automation—not general-purpose LLMs deployed without rigorous evaluation." The authors argue that "today's AI systems should not produce paper reviews" and that "current AI reviewers fail both conditions" (preservation of review diversity and resistance to gaming). They propose that "addressing the peer review crisis requires a science of peer review automation, including rigorous evaluation of specific tools for specific tasks, not wholesale deployment of general-purpose LLMs."

Risk of bias

Detection accuracy of AI-generated reviews relying on third-party classification labels (Emi 2025) with potential misclassification errors; Simulation homogeneity bias: use of single fixed prompt across models may understate diversity compared to real-world varied prompts; Self-preference bias in LLMs: GPT reviewers show larger score increases than Claude when GPT is the launderer (Panickssery et al., 2024); Sample selection bias: 60 randomly sampled papers may not represent full diversity of paper types and quality levels at ICLR; Embedding metric limitations: cosine similarity on text embeddings may not capture argumentative or evaluative diversity; AI-generation label classification errors from Emi (2025); Self-preference bias in LLMs (documented with GPT reviewers showing larger score increases than Claude); Experimental homogeneity in simulation (single fixed prompt may underestimate diversity); Sample representativeness: 60 papers may not represent full diversity of paper types; Detection limitations: embedding-based metrics may not capture argumentative diversity; AI-generation label classification errors from Emi (2025) - though validated via independent checks; Self-preference bias in LLMs (GPT reviewers showed larger score increases than Claude); Experimental homogeneity in simulation (single fixed prompt, two models only); Embedding-based metrics may not capture true argumentative diversity; Sample limited to ICLR venue and 60 papers - generalizability concerns; Selection bias in paper sampling - randomly selected but may not represent full distribution

Limitations

  • "Our AI reviewer simulations use only two models (GPT-5.1 and Claude Sonnet 4.5) with a single fixed prompt
  • In practice, researchers and conferences may use diverse prompts, temperatures, and model versions, which could yield more varied outputs
  • The high IntraSim thus partly reflects this experimental homogeneity rather than an intrinsic property of all possible AI reviewing setups." Additionally, "Our embedding-based similarity metrics capture linguistic and semantic patterns but do not directly measure diversity of viewpoints, arguments, or evaluative stances
  • Two reviews could be linguistically similar yet offer different critiques, or vice versa." The "simulation experiments use 60 randomly sampled ICLR papers
  • While sufficient to clearly show the two presented issues of gameability and non-diversity, this sample may not capture the full diversity of paper types and quality levels
  • Additionally, our findings are specific to one venue (ICLR) and may not generalize to conferences with different review norms or paper distributions."

Open questions raised

  • Need for adversarial robustness testing of AI systems in peer review before deployment
  • Need for empirical studies measuring diversity of reviewer opinions when AI assistance is provided
  • Need for user studies investigating how AI assistance affects reviewer behavior and whether humans catch AI errors
  • Need for large-scale surveys understanding what different stakeholders value about peer review
  • Need for research on how to enforce AI-usage policies in peer review
  • Need for metrics that directly assess argumentative diversity rather than linguistic similarity
Data: 75,800 ICLR 2026 reviews with AI-generation labels available at: iclr.pangram.com (from Emi 2025); ICLR 2026 reviews (75,800 reviews from 19,490 papers) with AI-generation labels available at iclr.pangram.com (Emi 2025); ICLR 2026 reviews with AI-generation labels available for download at: iclr.pangram.com (from Emi, 2025)Extracted from: pdfAgreement 59%

Explore related topics

Related papers