12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers

Lingyao Li, Junjie Xiong, Changjia Zhu, Runlong Yu, Chen Chen, Junyu Wang et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic benchmarking of 12 LLMs evaluated against 898 papers stratified from NeurIPS 2022, ICLR 2023, and ICLR 2025.

Sample

N = 898, 6 groups

Primary method

Descriptive statistics (mean gaps, percentage-point differences), Jensen-Shannon divergence for topic distribution comparison, readability metrics (Flesch-Kincaid, Gunning Fog), lexical diversity metrics (Type-Token Ratio), and controlled experimental comparison of clean vs. prompt-injected review conditions. Temperature set to 0 for reproducibility across LLM runs.

Main result

The study found that "LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary." Additionally, "Prompt injection remains highly effective" with "Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude: "While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks." They further state: "These findings indicate that LLMs are not substitutes for human judgment. Rather, they should be treated as potentially useful but unstable review-support tools whose behavior varies across models, tasks, and adversarial settings. Their integration into peer review requires calibration audits sensitive to model-family effects, defenses against hidden prompt injections, transparent reporting of model use, and clear policy boundaries on their role in acceptance decisions."

Risk of bias

Selection bias: papers sampled from only three top-tier AI venues (NeurIPS, ICLR) may not represent all research domains; Model selection bias: evaluation limited to mainstream LLMs from five providers, excluding smaller or specialized models; Temporal bias: data spans pre- and post-LLM availability periods but limited to 2022-2025 timeframe; Venue-specific bias: rating scale fixed at 10 points across venues despite different original scales; Selection bias: Papers sampled from OpenReview only; may not represent all conference submissions; Temporal bias: Data from 2022-2025; may not generalize to future LLM architectures; Model selection bias: Only 12 LLMs evaluated; selection criteria not fully transparent; Venue-specific bias: Limited to NeurIPS and ICLR; other conferences may differ; Rating scale heterogeneity: Different rating scales across venues (1-10 vs. discrete {1,3,5,6,8,10}) may introduce confounding; Prompt design bias: Shared system prompt may advantage certain model families over others; Annotation bias: Human reviews aggregated without weighting reviewer expertise; Reproducibility confound: Temperature set to 0 may not reflect production model behavior; Selection bias: Papers sampled from OpenReview may not represent all submission populations; Model selection bias: Evaluation limited to 12 mainstream LLMs from 5 providers, excludes smaller or specialized models; Venue bias: Focus on NeurIPS and ICLR may not generalize to other conferences or fields; Temporal bias: Different venues sampled in different years (2022, 2023, 2025), capturing different review norms

Limitations

  • "This research has several limitations that could be addressed in future research
  • First, our experiments primarily focus on mainstream LLMs (e.g., GPT, Gemini, Llama, Clau[de])" [text truncated in provided document]
  • The authors note that "Small track averages therefore mask large offsetting paper-level disagreements rather than reflecting LLM–human alignment" and that "stronger or newer models do not consistently show better paper-level alignment with human reviewers."

Open questions raised

  • Limited evaluation of non-mainstream LLMs and smaller model families
  • Need for broader venue coverage beyond NeurIPS and ICLR
  • Further investigation of defenses against prompt injection attacks
  • Long-term impacts of LLM-assisted review on peer review dynamics
  • Human-LLM collaboration models for peer review
  • Limited evaluation of non-mainstream LLMs and open-source alternatives beyond those tested
Data: Papers and reviews sourced from OpenReview (OpenReview, 2024), a widely adopted open-access platform hosting submissions and peer reviews for major computer science conferences. The dataset contains 898 papers stratified from NeurIPS 2022, ICLR 2023, and ICLR 2025.; OpenReview platform (OpenReview, 2024) used as primary data source; no indication that processed dataset is publicly available; OpenReview (OpenReview, 2024) - public platform with submissions and peer reviews from NeurIPS 2022, ICLR 2023, and ICLR 2025Code: Not mentioned in document; Not explicitly mentioned in the provided textExtracted from: pdfAgreement 50%

Explore related topics

Related papers