LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
Lingyao Li, Junjie Xiong, Changjia Zhu, Runlong Yu, Chen Chen, Junyu Wang et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic benchmarking of 12 LLMs evaluated against 898 papers stratified from NeurIPS 2022, ICLR 2023, and ICLR 2025.
Sample
N = 898, 6 groups
Primary method
Descriptive statistics (mean gaps, percentage-point differences), Jensen-Shannon divergence for topic distribution comparison, readability metrics (Flesch-Kincaid, Gunning Fog), lexical diversity metrics (Type-Token Ratio), and controlled experimental comparison of clean vs. prompt-injected review conditions. Temperature set to 0 for reproducibility across LLM runs.
Main result
The study found that "LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary." Additionally, "Prompt injection remains highly effective" with "Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude: "While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks." They further state: "These findings indicate that LLMs are not substitutes for human judgment. Rather, they should be treated as potentially useful but unstable review-support tools whose behavior varies across models, tasks, and adversarial settings. Their integration into peer review requires calibration audits sensitive to model-family effects, defenses against hidden prompt injections, transparent reporting of model use, and clear policy boundaries on their role in acceptance decisions."
Risk of bias
Selection bias: papers sampled from only three top-tier AI venues (NeurIPS, ICLR) may not represent all research domains; Model selection bias: evaluation limited to mainstream LLMs from five providers, excluding smaller or specialized models; Temporal bias: data spans pre- and post-LLM availability periods but limited to 2022-2025 timeframe; Venue-specific bias: rating scale fixed at 10 points across venues despite different original scales; Selection bias: Papers sampled from OpenReview only; may not represent all conference submissions; Temporal bias: Data from 2022-2025; may not generalize to future LLM architectures; Model selection bias: Only 12 LLMs evaluated; selection criteria not fully transparent; Venue-specific bias: Limited to NeurIPS and ICLR; other conferences may differ; Rating scale heterogeneity: Different rating scales across venues (1-10 vs. discrete {1,3,5,6,8,10}) may introduce confounding; Prompt design bias: Shared system prompt may advantage certain model families over others; Annotation bias: Human reviews aggregated without weighting reviewer expertise; Reproducibility confound: Temperature set to 0 may not reflect production model behavior; Selection bias: Papers sampled from OpenReview may not represent all submission populations; Model selection bias: Evaluation limited to 12 mainstream LLMs from 5 providers, excludes smaller or specialized models; Venue bias: Focus on NeurIPS and ICLR may not generalize to other conferences or fields; Temporal bias: Different venues sampled in different years (2022, 2023, 2025), capturing different review norms
Limitations
- "This research has several limitations that could be addressed in future research
- First, our experiments primarily focus on mainstream LLMs (e.g., GPT, Gemini, Llama, Clau[de])" [text truncated in provided document]
- The authors note that "Small track averages therefore mask large offsetting paper-level disagreements rather than reflecting LLM–human alignment" and that "stronger or newer models do not consistently show better paper-level alignment with human reviewers."
Open questions raised
- Limited evaluation of non-mainstream LLMs and smaller model families
- Need for broader venue coverage beyond NeurIPS and ICLR
- Further investigation of defenses against prompt injection attacks
- Long-term impacts of LLM-assisted review on peer review dynamics
- Human-LLM collaboration models for peer review
- Limited evaluation of non-mainstream LLMs and open-source alternatives beyond those tested
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations