Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community
Lin Li (28817), Qi Zhang (28502), Xander Davies, Jianing Qiu, Yarin Gal · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical study using adversarial optimization attacks against AI peer review systems.
Sample
N = 100, 7 groups
Primary method
Wilcoxon rank-sum test (one-sided, p < 0.05) for comparing pre- and post-rephrasing review scores. Bootstrap estimation for 95% confidence intervals around mean rating improvements (ΔScore). Normalized Shannon entropy for rating consistency. Iterative optimization algorithm with K=5 iterations, N=4 samples per iteration (N=8 for Meaning-Preserving with GPT 5.4 Mini), and M=6 reviews per rephrasing.
Main result
The study found that "AI reviewers are highly vulnerable to superficial rephrasing of the manuscript abstract. Rewriting only the abstract — a small part of the full paper (around 3.5% of total tokens per paper) and one that does not alter the underlying experiments, analyses or conclusions — can substantially inflate AI review evaluations. Our strongest attack achieves an attack success rate of about 38%, increasing acceptance ratings by +1.31 for Gemini 3 Flash reviewers and by +0.88 for GPT 5.4 Mini reviewers." When the original AI review suggests 'reject', the success rate rises to more than 50%.
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-computational (adversarial testing of AI systems)
Author conclusions
The authors conclude that "AI tools should not be treated as neutral evaluators in high-stakes peer review without systematic robustness testing, transparent safeguards and careful human oversight." They further state that "AI-driven peer review has the potential to meaningfully support the scientific community, but only if security, robustness, and incentive alignment are treated as central design requirements rather than afterthoughts" and that "AI-assisted peer review should be deployed with caution, transparent safeguards and systematic robustness evaluation before it is relied upon in high-stakes editorial decisions."
Risk of bias
Selection bias: Papers were selected from specific venues (ICLR 2025, Agents4Science 2025, Nature Communications), which may not be representative of all scientific disciplines or publication venues; Model-specific bias: Evaluation focused on two LLM models (GPT 5.4 Mini and Gemini 3 Flash); results may not generalize to other reviewer models; Review prompt bias: Different prompt formulations (Naive, Balanced, Complex) may elicit different vulnerabilities; Stochasticity: Multiple reviews were sampled (8 per paper) to account for LLM output variability, but this introduces sampling variability; Semantic equivalence assessment: Semantic equivalence check relies on LLM evaluation, which may incorrectly filter or retain rephrases; Selection bias: Papers were selected from rejected submissions, which may not represent the full distribution of submitted papers; Model-specific bias: Results depend on specific LLM configurations (GPT 5.4 Mini, Gemini 3 Flash) and may not generalize to other models; Prompt bias: The 'Balanced' prompt used for optimization may not generalize to significantly different review instructions; Stochasticity bias: LLM outputs are stochastic; results may vary with different random seeds; Dataset bias: Corpus is dominated by AI and medicine papers (65/100 and 23/100 respectively); Selection bias: Papers were selected from specific venues (ICLR 2025, Agents4Science 2025, Nature Communications), which may not represent all scientific domains equally; Model-specific bias: Only two primary AI models tested (GPT 5.4 Mini, Gemini 3 Flash); findings may not generalize to other reviewer models; Prompt dependency: Evaluation used three prompt configurations but other formulations might show different vulnerabilities; Paper sampling bias: Majority from AI domain (65/100 papers), with medicine (23/100) as secondary, limiting cross-disciplinary generalizability
Limitations
- The authors note that "all review and rephrasing experiments were conducted on the main manuscript text only" and appendices were removed, which may limit generalizability
- Additionally, the study evaluates only two AI models as reviewers, and the fluency constraint is instantiated "not as a specialized scientific language model, but for convenience as a general-purpose large language model that has been trained on a mixture of scientific and non-scientific text," which may not perfectly capture scientific writing norms
- The authors also acknowledge resource-dependent asymmetry, stating that "users with greater computational resources can more reliably push ratings upward and gain an advantage over lower-resource peers."
Open questions raised
- Need for systematic robustness testing of AI-assisted peer review systems
- Development of transparent safeguards for AI-mediated evaluation
- Investigation of defense mechanisms against abstract-rephrasing attacks
- Study of how inflated AI reviews bias downstream human editorial decisions
- Analysis of resource-dependent asymmetries in AI review manipulation
- The authors identify that "the robustness to strategic manipulation remains poorly understood" for AI-mediated peer review systems. They note a gap in understanding "subtler and potentially more pervasive risk" compared to overt prompt injection attacks. The paper suggests future work should focus on "systematic robustness testing, transparent safeguards and careful human oversight" of AI-assisted peer review systems.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations