12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the Future

Sihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu, Yiling Ma, Kaiyan Zhang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic narrative review with taxonomy construction.

Main result

The survey identifies a transformative shift in peer review driven by large language models and agent-based systems. Key findings include: "Around 2023, models began to process full manuscripts and produce credible end-to-end reviews" and "Current systems go beyond single-prompt generation to adopt multi-agent designs that emulate panel workflows and reinforcement learning to align outputs with nuanced human preferences." The paper reveals that "evaluation must keep pace" with methodological advances, establishing a taxonomy of five evaluation paradigms: human-centric, reference-based, LLM-based, aspect-oriented, and task-based evaluations.

Reports effect sizes.

Research paradigm

Interpretive/Systematic Literature Review

Author conclusions

"The landscape of peer review process has undergone a transformative shift, driven by advances in large language models and latest paradigms. This survey has provided a comprehensive synthesis of state-of-the-art methods, spanning review generation as well as after-review tasks including rebuttal, meta-review, and revision. We also present a systematic taxonomy of evaluation methods, metrics and datasets. Despite remarkable improvements in generation methods, significant challenges remain, including robust content-level evaluation and extensive capabilities. As the field advances, future research must prioritize transparent and ethical deployment of automated reviews."

Risk of bias

Domain concentration bias: Majority of datasets and evaluation studies concentrated in NLP and ML venues, not generalizable to other fields; Publication bias: Survey limited to publicly accessible literature; proprietary systems and unpublished advances not represented; Temporal bias: Rapid progress may outdated analyses soon after publication; Publication bias: Review concentrates on publicly accessible literature, excluding proprietary systems and unpublished industrial advances; Domain bias: Majority of datasets and evaluation studies concentrated in NLP and machine learning, limiting generalizability; Recency bias: Rapid pace of LLM advances may make portions of survey outdated soon after publication; Selection bias: Limited coverage of non-NLP/ML scientific domains; Domain concentration bias: majority of datasets focus on NLP and ML venues only; Publication bias: survey limited to publicly accessible literature, proprietary systems underrepresented; Temporal bias: rapid pace of LLM development may render findings outdated quickly; LLM judge biases: position bias, verbosity bias, style preference in LLM-based evaluations; Affiliation bias: studies show LLMs can exhibit favoritism toward highly-ranked institutions in single-blinded settings

Limitations

  • "While this survey strives to provide a comprehensive and up-to-date overview of the peer review landscape, several limitations remain
  • First, the rapid pace of progress in large language models, agent-based systems, and evaluation protocols means that new methodologies and benchmarks may emerge soon after publication, potentially outdating some of our analyses
  • Second, the majority of available datasets and evaluation studies are concentrated in NLP and machine learning domains, limiting the generalizability of our findings to other scientific fields
  • Finally, our survey is based on publicly accessible literature and resources
  • proprietary systems or unpublished industrial advances may not be adequately represented."

Open questions raised

  • Novelty Evaluation: LLMs remain weak at judging true novelty and distinguishing incremental from groundbreaking work
  • Automated Evaluation: Review quality remains hard to quantify; better content-aware evaluation methods needed
  • Beyond NLP and AI Domains: Most datasets concentrated in NLP/ML; broader disciplinary coverage needed
  • Beyond Review Generation: Automation of rebuttal, meta-review, and paper revision tasks remains underexplored
  • Multimodal Review Tasks: Current systems operate on text only; need for systems analyzing figures, tables, and code
  • Ethical and Transparent Deployment: Need for clear guidelines on ethical use and mandatory disclosure of LLM usage
Data: PEERREAD [Kang et al., 2018]; NLPEER [Dycke et al., 2023a]; ReviewMT [Tan et al., 2024]; MOPRD [Lin et al., 2022]; MRED [Shen et al., 2022]; PEERSUM [Li et al., 2023]; ORSUM [Zeng et al., 2024]; ARIES [D'Arcy et al., 2023]; CASIMIR [Jourdan et al., 2024]; ARXIVEDITS [Jiang et al., 2022]; ReviewCritique [Du et al., 2024]; SUBSTANREVIEW [Guo et al., 2023]; SchNovel [Lin et al., 2024]; NovBench [Wu et al., 2026b]; RR-MCQ [Zhou et al., 2024a]; Re2 [Zhang et al., 2025a]; REVIEWEVAL [Garg et al., 2025]; PEERREAD (Kang et al., 2018); NLPEER (Dycke et al., 2023a); ReviewMT (Tan et al., 2024); MOPRD (Lin et al., 2022); MRED (Shen et al., 2022); PEERSUM (Li et al., 2023); ORSUM (Zeng et al., 2024); ARIES (D'Arcy et al., 2023); ARXIVEDITS (Jiang et al., 2022); ReviewCritique (Du et al., 2024); SchNovel (Lin et al., 2024); NovBench (Wu et al., 2026b); SUBSTANREVIEW (Guo et al., 2023); RR-MCQ benchmark (Zhou et al., 2024a); Re2 (Zhang et al., 2025a); REVIEWEVAL (Garg et al., 2025)Extracted from: pdfAgreement 76%

Explore related topics

Related papers