CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
Hexuan Deng, Xiaopeng Ke, Yichen Li, Ruina Hu, Dehao Huang, Derek F. Wong et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational benchmark construction and evaluation.
Main result
The study found that "AI reviewers remain limited in correctness and are prone to hallucinations, and highlights reasoning models as more effective reviewers." Analysis reveals that "a single reviewer covers only 3.03 out of 5 top-level categories and 5.10 out of 23 subcategories on average," while "after combining all reviewers, a paper covers 3.98 categories out of 5 and 9.23 subcategories out of 23 on average." Additionally, "over 13% of submissions exhibit a score gap of 4 or more between the highest and lowest overall scores," and "in 22.13% of papers and 7.63% of reviews, incorrectness is identified via inter-reviewer conflicts."
Research paradigm
Empirical-computational; pragmatic benchmark construction combining systematic annotation with computational evaluation
Author conclusions
The authors conclude that "CoCoReviewBench enables category-level evaluation and detects erroneous reference comments to mitigate the limitations of human references." They further state: "Analysis shows that LLMs lag behind humans in correctness, are prone to hallucinations, and that reasoning models better substantiate their reviews." The study emphasizes that their benchmark "provides a feasible path to improve the completeness and correctness of references" and that findings "highlights limitations of human reviews as references for peer review evaluation and provides a feasible path to improve the completeness and correctness of references."
Risk of bias
Selection bias: Papers sampled with stratified sampling by score - may not represent full population; Meta-review bias: Meta-reviews may have inherent biases favoring reviewers over authors in negative decisions; Annotator bias: PhD student annotators may have disciplinary biases; Temporal bias: Data from specific years (2017-2025) may not generalize; Model bias: LLM-based classification and conflict detection may introduce systematic errors; Language/domain bias: Limited to ICLR and NeurIPS - primarily English ML/AI papers; Meta-review bias: annotators favor reviewers over authors when meta-review is negative; Acceptance status bias: lower annotation accuracy (50.03%) for rejected papers vs. accepted papers (79.76%); LLM-based classification may introduce hallucinations in segmentation and categorization steps; Stratified sampling by score may not represent all submission types equally; Selection of 'highest-scoring model' at each step could introduce bias toward particular LLM behaviors; Selection bias: Papers sampled uniformly across years (300 per year) with stratified sampling by score segments, which may not represent full distribution of submissions; Annotation bias: Human annotators' tendency to favor reviewers over authors when meta-reviews are negative; Model bias: LLM-based classification and conflict detection may inherit biases from training data; Temporal bias: Different annotation accuracy between accepted (79.76%) and rejected papers (50.03%); Meta-review bias: Coarse adjudication signal from meta-reviews may not perfectly reflect ground truth
Open questions raised
- Improved AI reviewer correctness and reduction of hallucinations
- Better handling of domain expertise in peer review evaluation
- Enhanced reviewer agreement mechanisms
- Methods for automated detection and correction of reviewer errors
- More sophisticated conflict resolution strategies beyond meta-review signals
- Better evaluation protocols beyond overlap metrics
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations