Review Arcade: On the Human Alignment and Gameability of LLM Reviews
Hans Ole Hatzel, Sebastian Steindl, Jan Strich · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical experiments on 984 papers from 2025 ACL Rolling Review (ARR).
Sample
N = 984, 2 groups
Primary method
Mean Absolute Error (MAE), Pearson's r correlation, paired t-tests with t/p-values, Cohen's d effect sizes, LLM judge evaluation using recall metrics (s_recall and w_recall), macro-averaging across accepted and rejected splits in Fisher-z space for correlation calculations
Main result
The study found that "LLM-human alignment varies substantially across prompts and models" and in the best-case scenario alignment is reasonable, but "this behavior does not consistently transfer to real-world conditions where the acceptance decisions are not known a priori." The authors also report that "in the constrained setup, 35% of papers improved after 10 rounds of edits, but this improvement also carried a risk of score regressions, with 22% of papers seeing a decrease in their score."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-Positivist
Author conclusions
The authors conclude: "Our results show that human-human correlation in review scores still surpasses the LLM-human alignment. Naively prompted LLMs are instable in their reviews and not yet generally reliable as peer-reviewers." They further state, "In this setup, it is feasible to use automated rewriting to push papers past the acceptance threshold in LLM-reliant peer-review." Most importantly, they warn: "We urge the community to employ extreme caution when approaching the subject of automated reviews. Given Goodhart's law, even when LLM reviews currently show decent alignment with human reviews, they might cease to be a good measure of submission quality."
Risk of bias
Selection bias: Rejected papers underrepresented in peer-review datasets (stratified subsampling used to address); Positivity bias: Existing research relies on very few or no rejected papers; Data imbalance: Rejected papers have fewer reviews per paper (1.1 vs 2.0 on average); Length bias: Rejected papers are substantially shorter than accepted papers (average 4,000-9,000 tokens vs 7,500+ tokens); Training data leakage: Uncertainty about whether LLMs have seen the test data during training; Selection bias: Dataset stratified to include ~33% rejected papers, which does not correspond to actual ARR acceptance rates; Underrepresentation of rejected papers in training data for human review correlation calculation; Data poisoning risk: LLMs may have seen ARR 2025 papers during training; Positivity bias: Previous ARR studies relied on few or no rejected papers; Imbalanced review counts: Accepted papers averaged 2.0 reviews; rejected papers averaged 1.1 reviews; PDF extraction variability: OCR model (olmOCR-2-7B-1025) used for paper conversion to Markdown; Positivity bias: Dataset weighted toward accepted papers in prior ARR research; authors address this through stratified subsampling; Representativeness bias: Rejected papers comprise roughly one-third of dataset, not matching actual ARR acceptance rates; Data contamination risk: LLMs may have been trained on papers in the test dataset; Limited human reviews for rejected papers (mean 1.1 reviews vs 2.0 for accepted papers); OCR extraction errors potentially affecting model performance assessment
Limitations
- The authors state: "Scores have the advantage of being easily quantifiable, but they also fail to account for many nuances in the utility of reviews." Additionally, "The best experiment to measure the effect of trying to game LLM reviews, is to review the edited submissions not only automatically, but also with humans
- This would allow to better understand if the edits are indeed improvements, or are simply superficial
- It is, however, virtually impossible to run such counterfactual reviews after the edits have been applied." Further, "Our dataset is limited in the number of reviews for rejected papers, leading to less reliable numbers, especially for the human-human correlation on the rejected split." The authors also note concerns about "Data Poisoning: It is possible that the LLMs we use have seen (part of) the data we test on during their training process."
Open questions raised
- Human evaluation of edited submissions: Need for counterfactual reviews with humans to determine if edits represent genuine improvements or superficial changes
- Cross-model generalization: Testing whether rephrasing attacks generalize to other models or to human reviewers
- Reviewer calibration: Inability to apply reviewer calibration methods due to limited number of reviews for rejected papers
- Holistic review assessment: Need to move beyond scores to evaluate full review content and nuance
- Comprehensive evaluation design: Call for extended evaluation of automated peer review with all strengths and weaknesses
- Need for counterfactual human reviews after edits to determine if improvements are genuine or superficial
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations