12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community

Lin Li (28817), Qi Zhang (28502), Xander Davies, Jianing Qiu, Yarin Gal · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical study using adversarial optimization attacks against AI peer review systems.

Sample

N = 100, 3 groups

Primary method

Wilcoxon rank-sum test (one-sided, p < 0.05) for comparing pre- and post-rephrasing review scores. Bootstrap estimation for 95% confidence intervals around mean rating improvements (ΔScore). Normalized Shannon entropy for rating consistency. Iterative optimization algorithm with K=5 iterations, N=4 samples per iteration (N=8 for Meaning-Preserving with GPT 5.4 Mini), and M=6 reviews per rephrasing.

Main result

The study found that "AI reviewers are highly vulnerable to superficial rephrasing of the manuscript abstract. Rewriting only the abstract — a small part of the full paper (around 3.5% of total tokens per paper) and one that does not alter the underlying experiments, analyses or conclusions — can substantially inflate AI review evaluations. Our strongest attack achieves an attack success rate of about 38%, increasing acceptance ratings by +1.31 for Gemini 3 Flash reviewers and by +0.88 for GPT 5.4 Mini reviewers." When the original AI review suggests 'reject', the success rate rises to more than 50%.

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-computational (adversarial testing of AI systems)

Author conclusions

The authors conclude that "AI tools should not be treated as neutral evaluators in high-stakes peer review without systematic robustness testing, transparent safeguards and careful human oversight." They further state that "AI-driven peer review has the potential to meaningfully support the scientific community, but only if security, robustness, and incentive alignment are treated as central design requirements rather than afterthoughts" and that "AI-assisted peer review should be deployed with caution, transparent safeguards and systematic robustness evaluation before it is relied upon in high-stakes editorial decisions."

Risk of bias

Selection bias: Papers were selected from specific venues (ICLR 2025, Agents4Science 2025, Nature Communications), which may not be representative of all scientific disciplines or publication venues; Model-specific bias: Evaluation focused on two LLM models (GPT 5.4 Mini and Gemini 3 Flash); results may not generalize to other reviewer models; Review prompt bias: Different prompt formulations (Naive, Balanced, Complex) may elicit different vulnerabilities; Stochasticity: Multiple reviews were sampled (8 per paper) to account for LLM output variability, but this introduces sampling variability; Semantic equivalence assessment: Semantic equivalence check relies on LLM evaluation, which may incorrectly filter or retain rephrases; Selection bias: Papers were selected from rejected submissions, which may not represent the full distribution of submitted papers; results may vary with different random seeds; Dataset bias: Corpus is dominated by AI and medicine papers (65/100 and 23/100 respectively); Prompt dependency: Evaluation used three prompt configurations but other formulations might show different vulnerabilities

Limitations

  • The authors note that "all review and rephrasing experiments were conducted on the main manuscript text only" and appendices were removed, which may limit generalizability
  • Additionally, the study evaluates only two AI models as reviewers, and the fluency constraint is instantiated "not as a specialized scientific language model, but for convenience as a general-purpose large language model that has been trained on a mixture of scientific and non-scientific text," which may not perfectly capture scientific writing norms
  • The authors also acknowledge resource-dependent asymmetry, stating that "users with greater computational resources can more reliably push ratings upward and gain an advantage over lower-resource peers."

Open questions raised

  • Need for systematic robustness testing of AI-assisted peer review systems
  • Development of transparent safeguards for AI-mediated evaluation
  • Investigation of defense mechanisms against abstract-rephrasing attacks
  • Study of how inflated AI reviews bias downstream human editorial decisions
  • Analysis of resource-dependent asymmetries in AI review manipulation
Data: 100 scientific papers from ICLR 2025 (40 rejected submissions); 40 AI-generated manuscripts from Agents4Science Conference 2025; 20 published/in-press articles from Nature Communications; Papers converted to standardized Markdown representation; 40 rejected papers from Agents4Science Conference 2025 (available via data release); Papers evaluated on LibriSpeech and VCTK datasets (mentioned in examples)Extracted from: pdf

Explore related topics

Related papers