On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale expert annotation study in which 45 domain scientists across Physical, Biological, and Health Sciences spent 469 hours rating 2,960 individual review items from human-written and AI-generated reviews of 82 Nature-family papers on three dimensions: correctness, significance, and sufficiency of evidence.
Sample
N = 2960, 7 groups
Primary method
Paired t-tests with Cohen's d effect sizes for binary metrics; Wilcoxon signed-rank test with rank-biserial correlation r for ordinal significance scores. All inferential tests used 95% bootstrap confidence intervals (10,000 paper-level resamples, percentile method). Inter-rater agreement measured via Gwet's AC1 (primary) and Cohen's κ. Generalized Linear Mixed Model (GLMM) with paper-level random intercepts for robustness analysis (Appendix C). Cluster-bootstrap confidence intervals for win-rate estimates. Item-level rates analyzed via Wilson 95% CI. Unit of analysis: papers (n=82) with within-paper correlation respected via pooling of item-level annotations.
Main result
On a composite of all three quality dimensions (correctness, significance, and evidence sufficiency), "a reviewing agent powered by GPT-5.2 scores above each paper's top-rated human reviewer (60.0% vs. 48.2%, p = 0.009)", while "all three AI reviewers (including Gemini 3.0 Pro and Claude Opus 4.5) exceed the lowest-rated human across every dimension." Additionally, "AI reviewers surface a distinct 26% of issues no human raises" and "exhibit 16 recurring weaknesses humans do not share, such as limited subfield knowledge, lack of long context management over multiple files, and overly critical stance on minor issues."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed methods (empirical-quantitative annotation study combined with qualitative analysis)
Author conclusions
"Overall, our results position current AI reviewers as complements to, not substitutes for, human reviewers." Further, "For AI reviewer developers, the priority is closing the correctness gap and improving criticism calibration. Concrete next steps include inducing an understanding of subfield-specific norms, embedding better long-context management into LLM agents, and calibrating to expert judgment on when criticism is warranted versus inflated." The authors conclude that "AI reviewers shouldn't be evaluated against human reviews as an implicit gold standard, but on the same per-axis standards (correctness, significance, and evidence sufficiency) that domain experts apply to human reviews."
Risk of bias
Selection bias: Papers included only if they had public peer reviews and pre-review versions available—non-random subset of Nature-family papers; Annotator expertise bias: Domain scientist annotators may have disciplinary biases or familiarity with particular methodologies; Model capability bias: Three frontier models tested (GPT-5.2, Claude Opus 4.5, Gemini 3.0 Pro) may not represent typical AI reviewer deployments; Prompt engineering bias: AI reviewers received specific structured prompts that may not reflect real-world usage; Attrition in cascading structure: Significance and evidence ratings only for correct items, reducing sample sizes and introducing selection bias in those dimensions; Selection bias: Papers were selected only if they met specific criteria (public peer review, publicly available pre-review version, subfield expert match), which may not be representative of all scientific papers. Annotation bias: While inter-annotator agreement was measured, the cascading structure (significance conditional on correctness, evidence conditional on both) means disagreement compounds across dimensions. The three-level significance scale showed only moderate agreement (Gwet's AC1 = 0.44). Potential expertise bias: Domain scientists recruiting from limited pool of 45 experts across 25 institutions may reflect particular research community norms.; Selection bias: Papers limited to Nature-family journals with transparent peer review policy and available pre-review versions (non-representative sample); Annotator expertise matching: Limited to papers with available domain scientist expertise (may exclude interdisciplinary or emerging fields); Inter-annotator agreement variability: Cohen's κ = 0.28-0.31 (fair) for some dimensions, indicating moderate disagreement despite Gwet's AC1 being higher; Single annotation for majority of papers: Only 27 of 82 papers (33%) doubly annotated, reducing reliability of singleton annotations; Model selection: Only three frontier LLM models tested; results may not generalize to other AI systems; Evaluation criteria dependency: Item-level evaluation depends on how review items are manually decomposed from free-text reviews; Temporal specificity: Papers published 2020-2025; AI reviewer behavior may change with model updates
Open questions raised
- The authors identify several open questions: (1) whether the correctness-significance tradeoff persists as models improve; (2) whether the observed patterns generalize beyond Nature-family papers; (3) what governance norms should accompany operational AI deployment; and (4) how best to integrate AI reviewers into venue-level review workflows. They note the need to close the correctness gap and improve criticism calibration through inducing understanding of subfield-specific norms, embedding better long-context management in LLM agents, and calibrating severity against field-specific norms.
- The authors identify several open questions: (1) whether the tradeoff between correctness and significance persists as models improve, (2) whether the patterns generalize beyond Nature-family papers, (3) what governance norms should accompany operational AI deployment, and (4) how AI reviewers should be integrated into venue-level review workflows. They also note that "existing evaluations of AI reviewers have focused on whether their verdicts match human verdicts (e.g., score alignment, acceptance prediction), which is insufficient to characterize their capabilities and limits." The paper addresses this gap by evaluating at the level of individual review items rather than aggregate scores.
- Open questions identified include: (1) whether the correctness-significance tradeoff persists as models improve, (2) whether the patterns generalize beyond Nature-family papers, and (3) what governance norms should accompany operational AI deployment. The authors note that "the experiments from this paper provides the empirical infrastructure to answer them, and the urgency only grows as AI reviewers move further into operational deployment."
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations