SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts
Yuan Xin, Yixuan Weng, Minjun Zhu, Ying Ling, Chengwei Qin, Michael Hahn et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative experimental evaluation with co-evolutionary adversarial training framework.
Sample
N = 1786, 5 groups
Primary method
Paired t-tests for comparing rating distributions; Spearman correlation for ranking quality assessment; threshold calibration protocol using top-33% quantile to match ground-truth acceptance rate; group normalization for advantage computation in GRPO; KL divergence estimation with unbiased per-token estimator; DPO preference optimization using log-sigmoid loss. Policy gradient methods (GRPO) with discrete reward function. Software: 8× H100 80G GPUs; vLLM for serving reward model; DeepSpeed ZeRO-3 for distributed training.
Main result
SafeReview substantially improves robustness compared to the undefended baseline: "SafeReview achieves superior ranking preservation: under Qwen3-4B attacks, SafeReview attains a Spearman correlation of 0.409 compared to Static DPO's 0.343 (+19.2%), with similar improvements under Qwen3-8B (0.396 vs 0.349, +13.5%) and Llama (0.392 vs 0.343, +14.3%)." The framework "significantly reduces the acceptance rate of adversarially manipulated papers under adaptive GRPO-style attacks, while achieving the highest Spearman correlation with ground-truth scores among all defense methods."
Reports effect sizes.
Research paradigm
Empirical computational research with experimental validation
Author conclusions
"SafeReview, a novel adversarial framework for defending LLM-based peer review systems against adversarial hidden prompts. By adapting the Co-evolutionary Adversarial Training paradigm to the unique challenges of scholarly evaluation, we established a co-evolutionary training process where attack and defense capabilities develop in tandem, ensuring robust protection against evolving threats. Our work has broader implications for the security of LLM-assisted academic evaluation. As these systems become increasingly prevalent in conferences and journals, ensuring their integrity is paramount to maintaining scholarly standards. SafeReview provides a foundational framework for this security, demonstrating that adversarial training can effectively harden review systems against manipulation while preserving their ability to provide constructive, evidence-based feedback."
Risk of bias
Training and testing on Computer Science venues only (NeurIPS, ICLR) - limits generalization across academic domains; Evaluation restricted to instruction-style prompt injections - may not capture other adversarial perturbation types; Distributional shift between training (NeurIPS 2024) and testing (ICLR 2024) - partially addresses overfitting but venues are similar; Group size G=8 in GRPO may introduce variance in reward estimation; Threshold calibration per condition rather than fixed threshold - could mask differences in calibration; Distributional shift between training and test sets mitigated but not fully addressed (NeurIPS 2024 training vs ICLR 2024 test); Limited to single defender architecture (DeepReviewer-14B) for main training; generalization tested but not guaranteed across all model families; Attack generation conditioned on taxonomy of known attack vectors may not capture novel adversarial strategies; Preference data construction uses binary clean/attacked distinction which may not reflect realistic gradations of adversarial manipulation; Closed-source model evaluation (Gemini, GPT-5.4) is zero-shot transfer only; defense mechanism not directly adapted for these systems; Training-test distributional shift: training on NeurIPS 2024 and testing on ICLR 2024 reduces but does not eliminate conference-specific bias; Limited domain coverage: evaluation restricted to Computer Science venues; Attack threat model specificity: evaluation limited to instruction-style prompt injections, not other semantic perturbations; Closed-source reviewer evaluation: transferability tests use proprietary models (Gemini, GPT-5.4) without access to internal mechanisms
Limitations
- "Future work should explore extending this framework to multi-modal submissions and investigate the transferability of attacks across a broader set of reviewer models
- While we have demonstrated that GRPO-trained attackers transfer effectively to closed-source reviewers (Gemini, GPT-5.4), adapting the defence mechanism itself for proprietary API-based reviewers, where fine-tuning is infeasible, remains open—potentially through prompt-based defence strategies or output filtering
- Furthermore, empirical validation is restricted to Computer Science venues (e.g., NeurIPS) and to specific instruction-style prompt injections, limiting the assessment of generalizability across diverse academic domains and robustness against more subtle, non-instruction-based semantic perturbations."
Open questions raised
- Extension to multi-modal submissions
- Broader transferability testing across diverse reviewer model families
- Defense mechanisms for proprietary API-based reviewers where fine-tuning is infeasible
- Generalization across diverse academic domains beyond Computer Science
- Robustness against subtle, non-instruction-based semantic perturbations
- Broader transferability of attacks across reviewer model families
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations