12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future

Sihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu, Yiling Ma, Kaiyan Zhang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic narrative review with taxonomy construction.

Main result

The survey identifies a transformative shift in peer review driven by large language models and agent-based systems. Key findings include: "Around 2023, models began to process full manuscripts and produce credible end-to-end reviews" and "Current systems go beyond single-prompt generation to adopt multi-agent designs that emulate panel workflows and reinforcement learning to align outputs with nuanced human preferences." The paper reveals that "evaluation must keep pace" with methodological advances, establishing a taxonomy of five evaluation paradigms: human-centric, reference-based, LLM-based, aspect-oriented, and task-based evaluations.

Reports effect sizes.

Research paradigm

Interpretive/Systematic Literature Review

Author conclusions

"The landscape of peer review process has undergone a transformative shift, driven by advances in large language models and latest paradigms. This survey has provided a comprehensive synthesis of state-of-the-art methods, spanning review generation as well as after-review tasks including rebuttal, meta-review, and revision. We also present a systematic taxonomy of evaluation methods, metrics and datasets. Despite remarkable improvements in generation methods, significant challenges remain, including robust content-level evaluation and extensive capabilities. As the field advances, future research must prioritize transparent and ethical deployment of automated reviews."

Risk of bias

Domain concentration bias: Majority of datasets and evaluation studies concentrated in NLP and ML venues, not generalizable to other fields; Publication bias: Survey limited to publicly accessible literature; proprietary systems and unpublished advances not represented; Temporal bias: Rapid progress may outdated analyses soon after publication; Selection bias: Limited coverage of non-NLP/ML scientific domains; Temporal bias: rapid pace of LLM development may render findings outdated quickly; LLM judge biases: position bias, verbosity bias, style preference in LLM-based evaluations; Affiliation bias: studies show LLMs can exhibit favoritism toward highly-ranked institutions in single-blinded settings

Limitations

  • "While this survey strives to provide a comprehensive and up-to-date overview of the peer review landscape, several limitations remain
  • First, the rapid pace of progress in large language models, agent-based systems, and evaluation protocols means that new methodologies and benchmarks may emerge soon after publication, potentially outdating some of our analyses
  • Second, the majority of available datasets and evaluation studies are concentrated in NLP and machine learning domains, limiting the generalizability of our findings to other scientific fields
  • Finally, our survey is based on publicly accessible literature and resources
  • proprietary systems or unpublished industrial advances may not be adequately represented."

Open questions raised

  • Novelty Evaluation: LLMs remain weak at judging true novelty and distinguishing incremental from groundbreaking work
  • Automated Evaluation: Review quality remains hard to quantify
  • better content-aware evaluation methods needed
  • Beyond NLP and AI Domains: Most datasets concentrated in NLP/ML
  • broader disciplinary coverage needed
  • Beyond Review Generation: Automation of rebuttal, meta-review, and paper revision tasks remains underexplored
Data: PEERREAD [Kang et al., 2018]; NLPEER [Dycke et al., 2023a]; ReviewMT [Tan et al., 2024]; MOPRD [Lin et al., 2022]; PEERSUM [Li et al., 2023]; NovBench [Wu et al., 2026b]; RR-MCQ [Zhou et al., 2024a]; Re2 [Zhang et al., 2025a]Extracted from: pdf

Explore related topics

Related papers