12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

Hans Ole Hatzel, Sebastian Steindl, Jan Strich · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Empirical experiments on 984 papers from 2025 ACL Rolling Review (ARR).

Sample

N = 984, 2 groups

Primary method

Mean Absolute Error (MAE), Pearson's r correlation, paired t-tests with t/p-values, Cohen's d effect sizes, LLM judge evaluation using recall metrics (s_recall and w_recall), macro-averaging across accepted and rejected splits in Fisher-z space for correlation calculations

Main result

The study found that "LLM-human alignment varies substantially across prompts and models" and in the best-case scenario alignment is reasonable, but "this behavior does not consistently transfer to real-world conditions where the acceptance decisions are not known a priori." The authors also report that "in the constrained setup, 35% of papers improved after 10 rounds of edits, but this improvement also carried a risk of score regressions, with 22% of papers seeing a decrease in their score."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-Positivist

Author conclusions

The authors conclude: "Our results show that human-human correlation in review scores still surpasses the LLM-human alignment. Naively prompted LLMs are instable in their reviews and not yet generally reliable as peer-reviewers." They further state, "In this setup, it is feasible to use automated rewriting to push papers past the acceptance threshold in LLM-reliant peer-review." Most importantly, they warn: "We urge the community to employ extreme caution when approaching the subject of automated reviews. Given Goodhart's law, even when LLM reviews currently show decent alignment with human reviews, they might cease to be a good measure of submission quality."

Risk of bias

Selection bias: Rejected papers underrepresented in peer-review datasets (stratified subsampling used to address); Positivity bias: Existing research relies on very few or no rejected papers; Data imbalance: Rejected papers have fewer reviews per paper (1.1 vs 2.0 on average); Length bias: Rejected papers are substantially shorter than accepted papers (average 4,000-9,000 tokens vs 7,500+ tokens); Training data leakage: Uncertainty about whether LLMs have seen the test data during training; Selection bias: Dataset stratified to include ~33% rejected papers, which does not correspond to actual ARR acceptance rates; Underrepresentation of rejected papers in training data for human review correlation calculation; Data poisoning risk: LLMs may have seen ARR 2025 papers during training; Positivity bias: Previous ARR studies relied on few or no rejected papers; Imbalanced review counts: Accepted papers averaged 2.0 reviews; rejected papers averaged 1.1 reviews; PDF extraction variability: OCR model (olmOCR-2-7B-1025) used for paper conversion to Markdown; Positivity bias: Dataset weighted toward accepted papers in prior ARR research; authors address this through stratified subsampling; Representativeness bias: Rejected papers comprise roughly one-third of dataset, not matching actual ARR acceptance rates; Data contamination risk: LLMs may have been trained on papers in the test dataset; Limited human reviews for rejected papers (mean 1.1 reviews vs 2.0 for accepted papers); OCR extraction errors potentially affecting model performance assessment

Limitations

  • The authors state: "Scores have the advantage of being easily quantifiable, but they also fail to account for many nuances in the utility of reviews." Additionally, "The best experiment to measure the effect of trying to game LLM reviews, is to review the edited submissions not only automatically, but also with humans
  • This would allow to better understand if the edits are indeed improvements, or are simply superficial
  • It is, however, virtually impossible to run such counterfactual reviews after the edits have been applied." Further, "Our dataset is limited in the number of reviews for rejected papers, leading to less reliable numbers, especially for the human-human correlation on the rejected split." The authors also note concerns about "Data Poisoning: It is possible that the LLMs we use have seen (part of) the data we test on during their training process."

Open questions raised

  • Human evaluation of edited submissions: Need for counterfactual reviews with humans to determine if edits represent genuine improvements or superficial changes
  • Cross-model generalization: Testing whether rephrasing attacks generalize to other models or to human reviewers
  • Reviewer calibration: Inability to apply reviewer calibration methods due to limited number of reviews for rejected papers
  • Holistic review assessment: Need to move beyond scores to evaluate full review content and nuance
  • Comprehensive evaluation design: Call for extended evaluation of automated peer review with all strengths and weaknesses
  • Need for counterfactual human reviews after edits to determine if improvements are genuine or superficial
Data: NLPeer dataset (Dycke et al., 2023); ACL Rolling Review (ARR) submissions from ACL 2025; ACL Rolling Review (ARR) 2025 submissions (984 papers, subset of NLPeer dataset); NLPeer dataset (Dycke et al., 2023) - used for stratified subsampling; NLPeer dataset (Dycke et al., 2023) - used for stratified subsampling to create 984-paper dataset; 2025 ACL Rolling Review (ARR) submissions - 984 papers with human reviewsCode: GitHub Repository (mentioned: "We publish our code." with footnote reference but specific URL not provided in text); GitHub Repository (mentioned but URL not provided in text); GitHub Repository (mentioned as published but specific URL not provided in text: "We publish our code.")Extracted from: pdfAgreement 56%

Explore related topics

Related papers