Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future
Sihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu, Yiling Ma, Kaiyan Zhang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic narrative review with taxonomy construction.
Main result
The survey identifies a transformative shift in peer review driven by large language models and agent-based systems. Key findings include: "Around 2023, models began to process full manuscripts and produce credible end-to-end reviews" and "Current systems go beyond single-prompt generation to adopt multi-agent designs that emulate panel workflows and reinforcement learning to align outputs with nuanced human preferences." The paper reveals that "evaluation must keep pace" with methodological advances, establishing a taxonomy of five evaluation paradigms: human-centric, reference-based, LLM-based, aspect-oriented, and task-based evaluations.
Reports effect sizes.
Research paradigm
Interpretive/Systematic Literature Review
Author conclusions
"The landscape of peer review process has undergone a transformative shift, driven by advances in large language models and latest paradigms. This survey has provided a comprehensive synthesis of state-of-the-art methods, spanning review generation as well as after-review tasks including rebuttal, meta-review, and revision. We also present a systematic taxonomy of evaluation methods, metrics and datasets. Despite remarkable improvements in generation methods, significant challenges remain, including robust content-level evaluation and extensive capabilities. As the field advances, future research must prioritize transparent and ethical deployment of automated reviews."
Risk of bias
Domain concentration bias: Majority of datasets and evaluation studies concentrated in NLP and ML venues, not generalizable to other fields; Publication bias: Survey limited to publicly accessible literature; proprietary systems and unpublished advances not represented; Temporal bias: Rapid progress may outdated analyses soon after publication; Publication bias: Review concentrates on publicly accessible literature, excluding proprietary systems and unpublished industrial advances; Domain bias: Majority of datasets and evaluation studies concentrated in NLP and machine learning, limiting generalizability; Recency bias: Rapid pace of LLM advances may make portions of survey outdated soon after publication; Selection bias: Limited coverage of non-NLP/ML scientific domains; Domain concentration bias: majority of datasets focus on NLP and ML venues only; Publication bias: survey limited to publicly accessible literature, proprietary systems underrepresented; Temporal bias: rapid pace of LLM development may render findings outdated quickly; LLM judge biases: position bias, verbosity bias, style preference in LLM-based evaluations; Affiliation bias: studies show LLMs can exhibit favoritism toward highly-ranked institutions in single-blinded settings
Limitations
- "While this survey strives to provide a comprehensive and up-to-date overview of the peer review landscape, several limitations remain
- First, the rapid pace of progress in large language models, agent-based systems, and evaluation protocols means that new methodologies and benchmarks may emerge soon after publication, potentially outdating some of our analyses
- Second, the majority of available datasets and evaluation studies are concentrated in NLP and machine learning domains, limiting the generalizability of our findings to other scientific fields
- Finally, our survey is based on publicly accessible literature and resources
- proprietary systems or unpublished industrial advances may not be adequately represented."
Open questions raised
- Novelty Evaluation: LLMs remain weak at judging true novelty and distinguishing incremental from groundbreaking work
- Automated Evaluation: Review quality remains hard to quantify; better content-aware evaluation methods needed
- Beyond NLP and AI Domains: Most datasets concentrated in NLP/ML; broader disciplinary coverage needed
- Beyond Review Generation: Automation of rebuttal, meta-review, and paper revision tasks remains underexplored
- Multimodal Review Tasks: Current systems operate on text only; need for systems analyzing figures, tables, and code
- Ethical and Transparent Deployment: Need for clear guidelines on ethical use and mandatory disclosure of LLM usage
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations