AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
Joydeep Biswas, Sheila Schoepp, Gautham Vasan, Anthony Opipari, Arthur Zhang, Zichao Hu et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale field deployment (22,977 papers) with: (1) multi-stage AI review system using frontier LLM (GPT-5) with specialized prompts, code interpreters, and web search tools; (2) large-scale voluntary survey (5,834 responses) of authors, program committee members, senior program committee members, and area chairs using 5-point Likert scale questionnaires and open-ended responses; (3) novel SPECS benchmark using synthetic perturbations of accepted papers to evaluate error detection across five criteria (Story, Presentation, Evaluations, Correctness, Significance); (4) qualitative thematic coding of 320 free-form survey responses..
Sample
N = 5834, 18 groups
Primary method
Mann-Whitney U tests used to assess statistical significance of differences in response distributions between AI and human reviews at α=0.01 level. Two-sided exact McNemar test used for SPECS benchmark comparisons. Five-point Likert scale responses analyzed using mean differences and distribution comparisons. Qualitative analysis used LLM-assisted taxonomy creation and classification of free-form responses.
Main result
The study found that "participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions." Specifically, "AI reviews were preferred to human reviews on six of the nine criteria" evaluated, with the largest advantages for AI reviews in "identifying technical errors (+0.67), raising previously unconsidered points (+0.61), suggesting improvements to presentation (+0.54) and research design (+0.49), and overall thoroughness (+0.48)." The system demonstrated operational feasibility, with "all reviews were generated in less than 24 hours" for "22,977 full-review papers" at a cost of "less than $1 per paper."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-pragmatist (mixed-methods: quantitative survey analysis + qualitative thematic coding + computational benchmarking)
Author conclusions
"The AAAI-26 AI Review Pilot Program demonstrated that AI-generated peer reviews are operationally feasible at conference scale, and are capable of generating reviews that are helpful to the authors and reviewers." The authors conclude that "state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research." They further state "the technical capabilities of this AI review system are already sufficient to usefully assist in scientific peer review in ways that the community finds helpful. Further study is needed to ascertain how best to integrate the complementary strengths of AI systems and human reviewers to improve the process of evaluating and advancing scientific research."
Risk of bias
Self-selection bias in survey participation (voluntary responses, n=5,834 from much larger population); Response bias: dissatisfied respondents may be over-represented in written feedback; Potential author bias: authors may have preference for AI reviews that are more uniformly thorough; Reviewer potential bias: experienced reviewers (SPC/AC) may rate reviews differently than regular PC members; Self-selection bias in voluntary survey participation (respondents may have stronger opinions about AI reviews); Potential measurement bias: respondents knew reviews were AI-generated (clearly labeled), which may have influenced ratings; Selection bias in benchmark dataset: papers required successful LaTeX compilation and matching arXiv sources, limiting representativeness; Possible response bias favoring novelty or experimental conditions; Self-selection bias in voluntary survey (5,834 of potentially much larger population); Selection bias in SPECS benchmark curation (only papers with compilable LaTeX sources from arXiv); Attrition: Authors and reviewers who did not respond to survey; Potential social desirability bias in survey responses regarding AI utility; Funding bias: OpenAI provided in-kind API credits and the study used OpenAI's GPT-5 model exclusively; Respondent role variation: Authors rated AI reviews more favorably than PC/SPC/ACs, introducing heterogeneous response patterns
Limitations
- The authors note that "evaluation remains challenging, however, because existing benchmarks and evaluation datasets measure only limited aspects of reviewing
- These aspects include specific error types, evaluation of structured outputs as opposed to unstructured review text, or similarity of scores to human reviews, rather than end-to-end review quality." Additionally, respondents identified that AI reviews had "errors in reading some equations and tables, difficulty in prioritizing the significance of issues (an area of ongoing research), and producing reviews that were longer than readers preferred." The qualitative analysis revealed that "reviews often failed to accurately assess the novelty, significance of contributions, and overall scientific impact of a paper" and showed "Shallow Contextual and Domain Understanding."
Open questions raised
- The authors identify that "despite substantial recent progress, a key question remained: could state-of-the-art AI systems generate technically meaningful and practically useful reviews in a live peer-review process at conference scale?" They also note that "Further study is needed to ascertain how best to integrate the complementary strengths of AI systems and human reviewers to improve the process of evaluating and advancing scientific research." Future work should address: (1) how to mitigate AI reviews' tendency to overemphasize minor issues, (2) improving handling of mathematical notation and tables, (3) better prioritization of significance of issues, (4) reducing review length, and (5) improving domain-specific understanding.
- How to best integrate complementary strengths of AI systems and human reviewers
- Improving AI system's ability to prioritize significance of issues (acknowledged as ongoing research area)
- Reducing verbosity in AI-generated reviews
- Improving accuracy in reading equations, tables, and mathematical notation
- Developing better domain-specific contextual understanding
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations