12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale expert annotation study in which 45 domain scientists across Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual review items from human-written and AI-generated reviews of 82 Nature-family papers on three dimensions: correctness, significance, and sufficiency of evidence.

Sample

N = 2960, 4 groups

Primary method

Paired t-tests with Cohen's d effect sizes for binary metrics; Wilcoxon signed-rank test with rank-biserial correlation r for ordinal significance scores. All inferential tests used 95% bootstrap confidence intervals (10,000 paper-level resamples, percentile method). Inter-rater agreement measured via Gwet's AC1 (primary) and Cohen's κ. Generalized Linear Mixed Model (GLMM) with paper-level random intercepts for robustness analysis (Appendix C). Cluster-bootstrap confidence intervals for win-rate estimates. Item-level rates analyzed via Wilson 95% CI. Unit of analysis: papers (n=82) with within-paper correlation respected via pooling of item-level annotations.

Main result

On a composite of all three quality dimensions (correctness, significance, and evidence sufficiency), "a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009)", while "all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension." Additionally, "AI reviewers surface a distinct 26% of issues no human raises" and "exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues."

Reports effect sizes and confidence intervals.

Research paradigm

Mixed methods (empirical-quantitative annotation study combined with qualitative analysis)

Author conclusions

"Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers." Further, "For AI reviewer developers, the priority is closing the correctness gap and improving criticism calibration. Concrete next steps include inducing an understanding of subfield-specific norms, embedding better long-context management into LLM agents, and calibrating to expert judgment on when criticism is warranted versus inflated." The authors conclude that "AI reviewers shouldn't be evaluated against human reviews as an implicit gold standard, but on the same per-axis standards (correctness, significance, and evidence sufficiency) that domain experts apply to human reviews."

Risk of bias

Selection bias: Papers included only if they had public peer reviews and pre-review versions available—non-random subset of Nature-family papers; Annotator expertise bias: Domain scientist annotators may have disciplinary biases or familiarity with particular methodologies; Model capability bias: Three frontier models tested (GPT-5.2, Claude Opus 4.5, Gemini 3.0 Pro) may not represent typical AI reviewer deployments; Prompt engineering bias: AI reviewers received specific structured prompts that may not reflect real-world usage; Attrition in cascading structure: Significance and evidence ratings only for correct items, reducing sample sizes and introducing selection bias in those dimensions; Annotator expertise matching: Limited to papers with available domain scientist expertise (may exclude interdisciplinary or emerging fields); Inter-annotator agreement variability: Cohen's κ = 0.28-0.31 (fair) for some dimensions, indicating moderate disagreement despite Gwet's AC1 being higher; Single annotation for majority of papers: Only 27 of 82 papers (33%) doubly annotated, reducing reliability of singleton annotations; results may not generalize to other AI systems; Evaluation criteria dependency: Item-level evaluation depends on how review items are manually decomposed from free-text reviews; Temporal specificity: Papers published 2020-2025; AI reviewer behavior may change with model updates

Open questions raised

  • The authors identify several open questions:
  • whether the correctness-significance tradeoff persists as models improve
  • whether the observed patterns generalize beyond Nature-family papers
  • what governance norms should accompany operational AI deployment
  • and (4) how best to integrate AI reviewers into venue-level review workflows. They note the need to close the correctness gap and improve criticism calibration through inducing understanding of subfield-specific norms, embedding better long-context management in LLM agents, and calibrating severity against field-specific norms.
Data: PeerReview Bench: a 78-paper benchmark dataset (released by authors); Expert annotation study data: 2,960 review items from 82 papers with expert ratings (availability not explicitly stated but implied as releasable); 82 Nature-family papers with peer reviews extracted from Research Square and Nature journalsCode: CMU PAPER REVIEWER: open-source AI reviewer platform (https://prometheus-eval.github.io/cmu-paper-reviewer/); OpenHands software-agent-sdk (used for deploying AI reviewers)Extracted from: pdf

Explore related topics

Related papers