12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions

Xiaomin Yang, Zhizhou Sha, Junbo Li, Jian Yu, Y W Sun, Matthew Zhao et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical adversarial attack study using closed-loop iterative optimization.

Sample

N = 500, 5 groups

Primary method

Wilcoxon signed-rank test (non-parametric test for score differences); Pairwise comparison with LLM judge using structured delta assessment (Δstrength, Δseverity on [-10, +10] scales); Directional aggregation using LLM synthesis rather than arithmetic mean to reduce noise from reviewer stochasticity; Direction gate with thresholds: τ_s=1.0 (strength improvement), τ_w=1.0 (weakness tolerance), τ_n=0.8 (net benefit); Composite selection score: SelectionScore = w_r·r̄_cand + w_c·Δ_content with weights w_r=0.8, w_c=0.2; Pearson and Spearman correlation for score calibration validation; Directional accuracy analysis for pairwise judge validation

Main result

The study found that "adversarial repackaging achieves a 75.1% attack success rate and a mean score gain of +1.21/10" across three mainstream AI reviewers. Key structural deficiencies include: "AI reviewers are easier to impress than to convince: highlighting strengths reliably increases perceived merit, while attempts to dissolve weaknesses frequently backfire." Additionally, "strategies that change how the reviewer interprets the paper, such as related-work repositioning and analytical discussion expansion, substantially outperform surface edits such as local polishing, table formatting, and algorithm boxes."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist; computational adversarial testing

Author conclusions

The authors conclude: "We propose resistance to presentation-only review gaming as a necessary condition for AI review automation and systematically test it through adversarial repackaging. Our experiments demonstrate that current AI reviewers fail this condition: with scientific content held entirely fixed, presentation-level edits alone suffice to raise review scores across multiple mainstream models and review templates." They further state: "This susceptibility has clear structural roots: strategies that reframe how the reviewer understands the paper are far more effective than those that improve its surface appearance, and attempts to dissolve specific criticisms frequently backfire. These findings indicate that the evaluation mechanisms of current AI reviewers can be systematically distorted through presentation-level manipulation."

Risk of bias

Selection bias: Dataset includes only papers with compilable LaTeX sources and moderate review scores, potentially excluding papers with poor formatting or very strong/very weak content; Reviewer stochasticity: Although mitigated through multiple independent reviews (N=3), AI models may exhibit biased responses to certain presentation styles; Model-specific bias: Testing limited to Claude and GPT families; generalization to other LLM families unknown; Pairwise judge bias: LLM judge (Claude Sonnet 4.5) used to assess directional changes may exhibit position bias, though order invariance tests showed high consistency (98-100%); Temperature setting confound: Different temperatures for reviewer (0.1 Claude, 1.0 GPT) and attacker (0.9 Claude, 1.0 GPT) models may affect comparability; Publication bias: Only unpublished arXiv papers included; published papers excluded due to training data contamination concerns; Limited model diversity: only three AI reviewer models tested (Claude Sonnet 4, Claude Sonnet 4.5, GPT-5-mini); Dataset contamination risk despite filtering: papers from arXiv may exist in LLM training data, though the authors implemented publication verification through Semantic Scholar; Selection bias in dataset construction: papers retained only if scoring in moderate range (mid-tier scores), excluding very weak or already-strong papers, affecting representativeness; Pairwise judge bias: all experiments use Claude Sonnet 4.5 as the pairwise judge, potentially introducing systematic bias in directional assessment; Attack agent configuration bias: matched-setting attacks use same model as reviewer but at different temperatures (0.9 for attacker vs 0.1 for reviewer in Claude models); Strategy attribution confounding: first-hit attribution does not control for strategy interaction effects; multiple strategies applied simultaneously in successful rounds; Temperature settings differ between models (Claude 0.1 vs GPT-5-mini 1.0), potentially affecting reproducibility; Attack agent uses higher temperature (0.9 for Claude, 1.0 for GPT) than reviewer (0.1 for Claude), introducing systematic differences; Selection bias in dataset construction: papers must have compilable LaTeX sources and fall within specific quality ranges

Limitations

  • The authors state: "Constrained by our compute budget, we have not yet tested additional models
  • However, the consistent vulnerability across all three configurations (and the positive cross-model transfer results, where attacks optimized against one model remain effective on another) suggests that the vulnerability reflects a structural property of current AI reviewers rather than a model-specific artifact." Additionally, "We observe a natural effectiveness ceiling: for papers whose weaknesses are grounded in concrete experimental gaps (e.g., single-dataset evaluation, absence of real-world validation), presentation optimization can still improve scores but the gains plateau around 5.0–5.5 rather than continuing to climb, with most improvement concentrated in the first few rounds
  • This indicates that the vulnerability is bounded: AI reviewers retain partial sensitivity to substantive shortcomings that presentation-only edits cannot fully override."

Open questions raised

  • Testing on additional AI reviewer models beyond Claude and GPT families
  • Broader deployment implications for AI-assisted peer review infrastructure
  • Development of countermeasures to defend AI reviewers against presentation-only gaming
  • Investigation of whether satisfying the presentation-only gaming resistance condition alone would justify safe AI reviewer deployment
  • Understanding effectiveness ceiling for papers with concrete experimental gaps
  • Investigating deployment recommendations and alternative explanations (mentioned as discussed in Appendix I but not detailed in main text)
Data: Contamination-free rolling benchmark of 500+ unpublished arXiv preprints with LaTeX sources and compiled PDFs (available through project website https://xyimatvoid.github.io/ARGAR-Site/); ICLR 2024–2025 historical peer review data from smallari/openreview-iclr-peer-reviews dataset (used for robustness validation); authors state "We release a contamination-free rolling dataset of recent unpublished arXiv preprints paired with their LATEX sources and PDFs using an automatic multi-stage filtering pipeline"; Dataset construction details provided in Appendix BCode: Project website with attack framework and benchmark: https://xyimatvoid.github.io/ARGAR-Site/Extracted from: pdf

Explore related topics

Related papers