12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Joydeep Biswas, Sheila Schoepp, Gautham Vasan, Anthony Opipari, Arthur Zhang, Zichao Hu et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale field deployment (22,977 papers) with: (1) multi-stage AI review system using frontier LLM (GPT-5) with specialized prompts, code interpreters, and web search tools; (2) large-scale voluntary survey (5,834 responses) of authors, program committee members, senior program committee members, and area chairs using 5-point Likert scale questionnaires and open-ended responses; (3) novel SPECS benchmark using synthetic perturbations of accepted papers to evaluate error detection across five criteria (Story, Presentation, Evaluations, Correctness, Significance); (4) qualitative thematic coding of 320 free-form survey responses..

Sample

N = 5834, 18 groups

Primary method

Mann-Whitney U tests used to assess statistical significance of differences in response distributions between AI and human reviews at α=0.01 level. Two-sided exact McNemar test used for SPECS benchmark comparisons. Five-point Likert scale responses analyzed using mean differences and distribution comparisons. Qualitative analysis used LLM-assisted taxonomy creation and classification of free-form responses.

Main result

The study found that "participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions." Specifically, "AI reviews were preferred to human reviews on six of the nine criteria" evaluated, with the largest advantages for AI reviews in "identifying technical errors (+0.67), raising previously unconsidered points (+0.61), suggesting improvements to presentation (+0.54) and research design (+0.49), and overall thoroughness (+0.48)." The system demonstrated operational feasibility, with "all reviews were generated in less than 24 hours" for "22,977 full-review papers" at a cost of "less than $1 per paper."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-pragmatist (mixed-methods: quantitative survey analysis + qualitative thematic coding + computational benchmarking)

Author conclusions

"The AAAI-26 AI Review Pilot Program demonstrated that AI-generated peer reviews are operationally feasible at conference scale, and are capable of generating reviews that are helpful to the authors and reviewers." The authors conclude that "state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research." They further state "the technical capabilities of this AI review system are already sufficient to usefully assist in scientific peer review in ways that the community finds helpful. Further study is needed to ascertain how best to integrate the complementary strengths of AI systems and human reviewers to improve the process of evaluating and advancing scientific research."

Risk of bias

Self-selection bias in survey participation (voluntary responses, n=5,834 from much larger population); Response bias: dissatisfied respondents may be over-represented in written feedback; Potential author bias: authors may have preference for AI reviews that are more uniformly thorough; Reviewer potential bias: experienced reviewers (SPC/AC) may rate reviews differently than regular PC members; Self-selection bias in voluntary survey participation (respondents may have stronger opinions about AI reviews); Potential measurement bias: respondents knew reviews were AI-generated (clearly labeled), which may have influenced ratings; Selection bias in benchmark dataset: papers required successful LaTeX compilation and matching arXiv sources, limiting representativeness; Possible response bias favoring novelty or experimental conditions; Self-selection bias in voluntary survey (5,834 of potentially much larger population); Selection bias in SPECS benchmark curation (only papers with compilable LaTeX sources from arXiv); Attrition: Authors and reviewers who did not respond to survey; Potential social desirability bias in survey responses regarding AI utility; Funding bias: OpenAI provided in-kind API credits and the study used OpenAI's GPT-5 model exclusively; Respondent role variation: Authors rated AI reviews more favorably than PC/SPC/ACs, introducing heterogeneous response patterns

Limitations

  • The authors note that "evaluation remains challenging, however, because existing benchmarks and evaluation datasets measure only limited aspects of reviewing
  • These aspects include specific error types, evaluation of structured outputs as opposed to unstructured review text, or similarity of scores to human reviews, rather than end-to-end review quality." Additionally, respondents identified that AI reviews had "errors in reading some equations and tables, difficulty in prioritizing the significance of issues (an area of ongoing research), and producing reviews that were longer than readers preferred." The qualitative analysis revealed that "reviews often failed to accurately assess the novelty, significance of contributions, and overall scientific impact of a paper" and showed "Shallow Contextual and Domain Understanding."

Open questions raised

  • The authors identify that "despite substantial recent progress, a key question remained: could state-of-the-art AI systems generate technically meaningful and practically useful reviews in a live peer-review process at conference scale?" They also note that "Further study is needed to ascertain how best to integrate the complementary strengths of AI systems and human reviewers to improve the process of evaluating and advancing scientific research." Future work should address: (1) how to mitigate AI reviews' tendency to overemphasize minor issues, (2) improving handling of mathematical notation and tables, (3) better prioritization of significance of issues, (4) reducing review length, and (5) improving domain-specific understanding.
  • How to best integrate complementary strengths of AI systems and human reviewers
  • Improving AI system's ability to prioritize significance of issues (acknowledged as ongoing research area)
  • Reducing verbosity in AI-generated reviews
  • Improving accuracy in reading equations, tables, and mathematical notation
  • Developing better domain-specific contextual understanding
Data: SPECS benchmark dataset implied to be generated from AAAI-25 proceedings papers with arXiv LaTeX sources. Specific dataset URLs or availability statements are not explicitly provided in the paper. The paper states the curation process "can be repeated to generate a new dataset, e.g., for a different venue" but does not state the current dataset is publicly available.Code: No code repositories are explicitly mentioned as available. The paper references "olmOCR [34]" for PDF-to-markdown conversion but does not provide a repository link for the AAAI-26 AI Review System code.Extracted from: pdfAgreement 51%

Explore related topics

Related papers