12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale expert annotation study in which 45 domain scientists across Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual review items from human-written and AI-generated reviews of 82 Nature-family papers on three dimensions: correctness, significance, and sufficiency of evidence.

Sample

N = 2960, 7 groups

Primary method

Paired t-tests with Cohen's d effect sizes for binary metrics; Wilcoxon signed-rank test with rank-biserial correlation r for ordinal significance scores. All inferential tests used 95% bootstrap confidence intervals (10,000 paper-level resamples, percentile method). Inter-rater agreement measured via Gwet's AC1 (primary) and Cohen's κ. Generalized Linear Mixed Model (GLMM) with paper-level random intercepts for robustness analysis (Appendix C). Cluster-bootstrap confidence intervals for win-rate estimates. Item-level rates analyzed via Wilson 95% CI. Unit of analysis: papers (n=82) with within-paper correlation respected via pooling of item-level annotations.

Main result

On a composite of all three quality dimensions (correctness, significance, and evidence sufficiency), "a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009)", while "all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension." Additionally, "AI reviewers surface a distinct 26% of issues no human raises" and "exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues."

Reports effect sizes and confidence intervals.

Research paradigm

Mixed methods (empirical-quantitative annotation study combined with qualitative analysis)

Author conclusions

"Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers." Further, "For AI reviewer developers, the priority is closing the correctness gap and improving criticism calibration. Concrete next steps include inducing an understanding of subfield-specific norms, embedding better long-context management into LLM agents, and calibrating to expert judgment on when criticism is warranted versus inflated." The authors conclude that "AI reviewers shouldn't be evaluated against human reviews as an implicit gold standard, but on the same per-axis standards (correctness, significance, and evidence sufficiency) that domain experts apply to human reviews."

Risk of bias

Selection bias: Papers included only if they had public peer reviews and pre-review versions available—non-random subset of Nature-family papers; Annotator expertise bias: Domain scientist annotators may have disciplinary biases or familiarity with particular methodologies; Model capability bias: Three frontier models tested (GPT-5.2, Claude Opus 4.5, Gemini 3.0 Pro) may not represent typical AI reviewer deployments; Prompt engineering bias: AI reviewers received specific structured prompts that may not reflect real-world usage; Attrition in cascading structure: Significance and evidence ratings only for correct items, reducing sample sizes and introducing selection bias in those dimensions; Selection bias: Papers were selected only if they met specific criteria (public peer review, publicly available pre-review version, subfield expert match), which may not be representative of all scientific papers. Annotation bias: While inter-annotator agreement was measured, the cascading structure (significance conditional on correctness, evidence conditional on both) means disagreement compounds across dimensions. The three-level significance scale showed only moderate agreement (Gwet's AC1 = 0.44). Potential expertise bias: Domain scientists recruiting from limited pool of 45 experts across 25 institutions may reflect particular research community norms.; Selection bias: Papers limited to Nature-family journals with transparent peer review policy and available pre-review versions (non-representative sample); Annotator expertise matching: Limited to papers with available domain scientist expertise (may exclude interdisciplinary or emerging fields); Inter-annotator agreement variability: Cohen's κ = 0.28-0.31 (fair) for some dimensions, indicating moderate disagreement despite Gwet's AC1 being higher; Single annotation for majority of papers: Only 27 of 82 papers (33%) doubly annotated, reducing reliability of singleton annotations; Model selection: Only three frontier LLM models tested; results may not generalize to other AI systems; Evaluation criteria dependency: Item-level evaluation depends on how review items are manually decomposed from free-text reviews; Temporal specificity: Papers published 2020-2025; AI reviewer behavior may change with model updates

Open questions raised

  • The authors identify several open questions: (1) whether the correctness-significance tradeoff persists as models improve; (2) whether the observed patterns generalize beyond Nature-family papers; (3) what governance norms should accompany operational AI deployment; and (4) how best to integrate AI reviewers into venue-level review workflows. They note the need to close the correctness gap and improve criticism calibration through inducing understanding of subfield-specific norms, embedding better long-context management in LLM agents, and calibrating severity against field-specific norms.
  • The authors identify several open questions: (1) whether the tradeoff between correctness and significance persists as models improve, (2) whether the patterns generalize beyond Nature-family papers, (3) what governance norms should accompany operational AI deployment, and (4) how AI reviewers should be integrated into venue-level review workflows. They also note that "existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits." The paper addresses this gap by evaluating at the level of individual review items rather than aggregate scores.
  • Open questions identified include: (1) whether the correctness-significance tradeoff persists as models improve, (2) whether the patterns generalize beyond Nature-family papers, and (3) what governance norms should accompany operational AI deployment. The authors note that "the experiments from this paper provides the empirical infrastructure to answer them, and the urgency only grows as AI reviewers move further into operational deployment."
Data: PeerReview Bench: a 78-paper benchmark dataset (released by authors); Expert annotation study data: 2,960 review items from 82 papers with expert ratings (availability not explicitly stated but implied as releasable); PEERREVIEW BENCH: A 78-paper benchmark dataset for automatically evaluating AI reviewers, released by the authors. The 82 papers analyzed are from Nature-family journals with publicly released peer reviews and publicly available pre-review versions on Research Square (https://www.researchsquare.com/).; PEERREVIEWBENCH: 78-paper benchmark dataset for evaluating AI reviewers (https://prometheus-eval.github.io/cmu-paper-reviewer/); 82 Nature-family papers with peer reviews extracted from Research Square and Nature journalsCode: CMU PAPER REVIEWER: open-source AI reviewer platform (https://prometheus-eval.github.io/cmu-paper-reviewer/); OpenHands software-agent-sdk (used for deploying AI reviewers); CMU PAPER REVIEWER: An open-source AI reviewer platform available at https://prometheus-eval.github.io/cmu-paper-reviewer/. The paper states: "We release the CMU PAPER REVIEWER, an open-source platform for authors, students, and researchers who want detailed feedback on a manuscript before submission, built on the pipeline employed in our expert annotation study."Extracted from: pdfAgreement 47%

Explore related topics

Related papers