12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Generative Reviewer Agents: Scalable Simulacra of Peer Review

Nicolas Bougie, Narimawa Watanabe · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.emnlp-industry.8

Methodology & findings

Study design

Multi-round empirical evaluation framework.

Sample

N = 3797, 18 groups

Primary method

Bradley-Terry model for ranking preferences; quadratic-weighted Cohen's κ for ordinal score alignment; Pearson correlation coefficients for criterion alignment; t-tests for model performance comparison; binary cross-entropy loss minimization for BT coefficient estimation; win matrix construction; logistic regression. Results averaged over 20 runs (reviews) or 3 seeds (other experiments). Standard deviations reported as measures of variability.

Main result

The study found that GAR achieves an F1 score of 0.66 in predicting paper acceptance, "significantly exceeding human reviewers' 0.49 score" (p < 0.002). Furthermore, "GAR leads with a score of 0.684, outperforming human reviewers (0.523)" in pairwise preference comparisons using the Bradley-Terry model. The framework demonstrates that "GAR generated reviews stem from their depth, resulting in high preferences" through retrieval of relevant community-level reviews.

Reports effect sizes and confidence intervals.

Research paradigm

computational empiricism with agent-based simulation

Author conclusions

The authors conclude: "This research marks a step towards improving scientific writing and research by offering cost-effective, in-depth, and on-demand reviews." They further state: "Our vision is not to replace human reviewers but to enhance the review process by supporting them with synthetic reviewers capable of managing the increasing volume of submissions and providing early, constructive feedback. This collaboration between agents and human experts has the potential to accelerate scientific progress, foster innovation, and reduce time-to-publication."

Risk of bias

LLM training data bias: GAR may inherit biases from underlying LLMs (GPT-4o, Llama, Mistral); Institutional bias: potential preference for well-established institutions reflected in training data; Novelty detection bias: system overestimates low-novelty papers and underestimates highly novel papers; Domain-specific bias: evaluation limited to machine learning conferences; generalizability unknown; Information leakage: evaluation papers may have been partially included in LLM pre-training corpora; Memory contamination: title-matching filter to avoid review reuse is 'not foolproof' for modified manuscripts; Training data bias in LLMs affecting reviewer behavior simulation; Potential information leakage from pre-training corpora (papers may be in LLM training data); Institutional bias not fully mitigated (paper affiliation effects); Limited expertise representation in novelty assessment; Incomplete title/content matching for bias detection (fuzzy matching not implemented); Overestimation of low-novelty papers and underestimation of high-novelty papers; Lack of diversity in conference domains (ML conferences only); LLM training data bias - GAR may inherit biases from LLM training corpora; Information leakage - papers may have been included in LLM pre-training datasets; Institutional bias - though analysis suggests GAR may reduce rather than amplify institutional bias; Paradigm-shifting work detection - systematic bias against novel/unconventional ideas; Domain bias - trained/evaluated exclusively on machine learning conferences; Memory module contamination - risk of exact or near-duplicate papers biasing retrieval

Limitations

  • The authors state: "First, our study primarily focuses on isolating and evaluating specific factors in the peer review process, such as reviewer dedication or expertise, instead of accounting for the inherent variability and arbitrariness that occur in real peer review scenarios." Additionally, "A limitation is GAR's difficulty in evaluating highly novel or paradigm-shifting work, as noted in Appendix A.4
  • The system may struggle to recognize contributions that deviate from established norms, potentially overlooking groundbreaking ideas." Furthermore, "our evaluation was conducted in the context of machine learning conferences
  • as such, the generalizability of our findings to other domains or conference communities remains to be established." The authors also note potential information leakage: "Since most frontier LLMs are trained on proprietary datasets, it remains difficult to ascertain whether evaluation papers may have been partially included in the models' training data."

Open questions raised

  • Difficulty evaluating highly novel or paradigm-shifting work; future work should focus on equipping synthetic reviewers with ability to recognize novelty at a more nuanced level, including leveraging knowledge graph structure at community level or using citation embeddings
  • Generalizability beyond machine learning conferences to other disciplines (mathematics, experimental sciences) with different review norms
  • Handling sensitive information and intellectual property safeguards for practical deployment
  • Advanced contamination detection strategies (fuzzy matching or content-based similarity filtering) to prevent reusing prior reviews or content
  • Locally trained models on carefully controlled corpora to address information leakage from pre-training
  • Difficulty recognizing genuinely groundbreaking or paradigm-shifting work
Data: ICLR 2023 (3,797 papers from OpenReview); ICLR 2022 (from OpenReview); NeurIPS 2023 (from OpenReview); ICLR 2023 dataset (3,797 papers from Openreview); ICLR 2022 dataset; NeurIPS 2023 dataset; ICLR 2022 dataset (via Openreview); NeurIPS 2023 dataset (via Beygelzimer et al., 2021)Code: Not mentioned in the paper; Not statedExtracted from: pdfAgreement 46%

Explore related topics

Related papers