Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review
X Zhao, Rana Muhammad Shahroz Khan, Zhen Xu, Zhen Tan, T T Chen · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
The paper presents a comprehensive computational benchmark study combining: (1) black-box prompt injection attacks on text, (2) white-box gradient-based perturbations using GCG for text and PGD/APGD/C&W for images, (3) systematic evaluation across multiple state-of-the-art LLMs/MLLMs using a dataset of 1136 papers from ICLR and F1000Research, and (4) evaluation of three defense strategies (LLM-as-Judge, trained classifiers, chunk-based embedding search).
Main result
The study demonstrates that "AI reviewers are pervasively vulnerable" to adversarial manipulation. Specifically, "Black-box prompt injections achieve up to an 80% Attack Success Rate (ASR) against powerful proprietary models, causing massive score inflation. Similarly, white-box visual attacks inflate scores by up to +14.11 points, confirming the insufficiency of text-only safeguards." Additionally, "Claude-sonnet-4.5 exhibits the highest fragility, achieving an Attack Success Rate (ASR) of 0.80 and a massive score inflation (14.14 points)."
Research paradigm
Positivist/Empiricist (systematic benchmark development and quantitative evaluation)
Author conclusions
The authors conclude that "PaperGuard establishes the foundational benchmark, protocols, and actionable defense necessary to pioneer trustworthy, attack-resilient AI-assisted scholarly reviewing." They emphasize that "By establishing this foundational benchmark, transparent protocols, and an actionable defense, PaperGuard provides the critical tools necessary to pioneer trustworthy AI-assisted scholarly reviewing." Furthermore, they state that "Our extensive experiments across state-of-the-art models confirm that AI reviewers are pervasively vulnerable" and that their "chunk-based embedding search defense effectively detects malicious injections while produces zero false positive case, making it a practical solution that avoids penalizing benign authors."
Risk of bias
Selection bias in dataset construction: papers deliberately constrained to rejected ICLR and F1000 submissions with strong rejection leans, which may not represent acceptance scenarios. Potential gradient masking concerns in white-box attacks requiring diverse attack ensemble (PGD, APGD, C&W) to ensure robust evaluation. Model capability artifacts: DeepSeek-R1-Distill-Llama-8B's lower ASR attributed to capability failure rather than security alignment (formatting instruction failures ~38% of cases).; Selection bias: Dataset deliberately constrained to rejected papers with strong rejection leans to focus on challenging cases; Model capability confounds: Smaller models may show lower ASR due to capability failure rather than robustness; Hyperparameter sensitivity: White-box attacks' effectiveness varies with optimization choices (learning rates, iteration counts); Dataset construction bias: Deliberately selected rejected papers with strong rejection leans, not representative of all submissions; Model selection bias: Evaluation focuses on state-of-the-art models; older/smaller models may have different vulnerabilities; Attack transfer bias: White-box attacks optimized on surrogate models may not fully represent real deployment threats; Evaluation context bias: Models evaluated on anonymized papers following double-blind review protocols, which may differ from actual deployment scenarios
Limitations
- The authors note that "the black-box injection setting is the most directly deployable threat: the adversary needs no access to the reviewer's internals and only crafts content that is processed as part of the submission
- The white-box setting does not assume that a real attacker obtains exact gradient access to the deployed reviewer
- rather, it serves as (i) a stress-test upper bound on model vulnerability and (ii) a way to optimize attacks on an open surrogate model and study their transfer to other reviewers." Additionally, the document embedding classifier faces "a primary, unavoidable limitation: the context length of most document embedders (e.g., GTE-large, E5) is far smaller than the full paper
- We are therefore forced to auto-truncate the input text T to fit the model's context window."
Open questions raised
- Gap 1: Existing robustness studies overwhelmingly focus on text-only attacks, neglecting the visual modality where core methodology and results are presented. Gap 2: AI-review safety is distinct from standard jailbreaking, requiring domain-specific targeted failures rather than general safety policy violations. Gap 3: No practical defenses exist for peer-review attacks, particularly for subtle domain-specific manipulations embedded within lengthy documents.
- The paper identifies three key gaps in prior work: (Gap 1) Existing robustness studies are overwhelmingly focused on text-only attacks, neglecting the visual modality where core methodology and results are presented. (Gap 2) AI-review safety is distinct from standard jailbreaking; peer-review attacks seek domain-specific targeted failures rather than general safety policy violations, requiring different attack and defense approaches. (Gap 3) No practical defenses exist for domain-specific peer-review manipulation threats.
- Existing robustness studies focus overwhelmingly on text-only attacks, neglecting visual modality where core methodology and results are presented
- AI-review safety literature lacks understanding of domain-specific, targeted failures distinct from standard jailbreaking
- No practical defenses exist for adversarial manipulation in peer-review systems
- Limited understanding of multimodal adversarial vulnerabilities of AI reviewers
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations