12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Hexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu, Dehao Huang, Derek F. Wong et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational benchmark construction and evaluation.

Main result

The study found that "AI reviewers remain limited in correctness and are prone to hallucinations, and highlights reasoning models as more effective reviewers." Analysis reveals that "a single reviewer covers only 3.03 out of 5 top-level categories and 5.10 out of 23 subcategories on average," while "after combining all reviewers, a paper covers 3.98 categories out of 5 and 9.23 subcategories out of 23 on average." Additionally, "over 13% of submissions exhibit a score gap of 4 or more between the highest and lowest overall scores," and "in 22.13% of papers and 7.63% of reviews, incorrectness is identified via inter-reviewer conflicts."

Research paradigm

Empirical-computational; pragmatic benchmark construction combining systematic annotation with computational evaluation

Author conclusions

The authors conclude that "CoCoReviewBench enables category-level evaluation and detects erroneous reference comments to mitigate the limitations of human references." They further state: "Analysis shows that LLMs lag behind humans in correctness, are prone to hallucinations, and that reasoning models better substantiate their reviews." The study emphasizes that their benchmark "provides a feasible path to improve the completeness and correctness of references" and that findings "highlights limitations of human reviews as references for peer review evaluation and provides a feasible path to improve the completeness and correctness of references."

Risk of bias

Selection bias: Papers sampled with stratified sampling by score - may not represent full population; Meta-review bias: Meta-reviews may have inherent biases favoring reviewers over authors in negative decisions; Annotator bias: PhD student annotators may have disciplinary biases; Temporal bias: Data from specific years (2017-2025) may not generalize; Model bias: LLM-based classification and conflict detection may introduce systematic errors; Language/domain bias: Limited to ICLR and NeurIPS - primarily English ML/AI papers; Meta-review bias: annotators favor reviewers over authors when meta-review is negative; Acceptance status bias: lower annotation accuracy (50.03%) for rejected papers vs. accepted papers (79.76%); LLM-based classification may introduce hallucinations in segmentation and categorization steps; Stratified sampling by score may not represent all submission types equally; Selection of 'highest-scoring model' at each step could introduce bias toward particular LLM behaviors; Selection bias: Papers sampled uniformly across years (300 per year) with stratified sampling by score segments, which may not represent full distribution of submissions; Annotation bias: Human annotators' tendency to favor reviewers over authors when meta-reviews are negative; Model bias: LLM-based classification and conflict detection may inherit biases from training data; Temporal bias: Different annotation accuracy between accepted (79.76%) and rejected papers (50.03%); Meta-review bias: Coarse adjudication signal from meta-reviews may not perfectly reflect ground truth

Open questions raised

  • Improved AI reviewer correctness and reduction of hallucinations
  • Better handling of domain expertise in peer review evaluation
  • Enhanced reviewer agreement mechanisms
  • Methods for automated detection and correction of reviewer errors
  • More sophisticated conflict resolution strategies beyond meta-review signals
  • Better evaluation protocols beyond overlap metrics
Data: CoCoReviewBench: 3,900 papers with structured annotations from ICLR and NeurIPS; Available at: https://github.com/hexuandeng/CoCoReviewBench; CoCoReviewBench: 3,900 papers from ICLR (2017-2025) and NeurIPS (2021-2024) with structured annotations; available at https://github.com/hexuandeng/CoCoReviewBench; CoCoReviewBench dataset with 3,900 papers, 14.1k review comments, 134.8k opinions, 115.9k opinion clusters, and 108.6k correct opinions available at https://github.com/hexuandeng/CoCoReviewBenchCode: https://github.com/hexuandeng/CoCoReviewBench; https://github.com/hexuandeng/CoCoReviewBench (main benchmark and models)Extracted from: pdfAgreement 41%

Explore related topics

Related papers