12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time

Ilia Kuznetsov, Rohan Nayak, Alla Rozovskaya, Iryna Gurevych · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational measurement and comparative analysis framework.

Sample

N = 15683, 4 groups

Primary method

Bootstrap hypothesis testing (10,000 iterations, 99% confidence intervals, Bonferroni correction for multiple comparisons per venue). Spearman's rank correlation (ρ) for measuring relationships between measurements. Normalization of count-based and length measurements to [0,1] scale. Sentence-level alignment using all-MiniLM-L6-v2 embeddings with 0.8 similarity threshold. Median comparison across reviewing campaigns as primary descriptive statistic. Human inter-rater agreement assessed via manual annotation and accuracy calculation.

Main result

Contradicting the popular narrative, the study found that "the median review quality at computer science conferences shows no consistent decline" over time. Specifically, the authors report "no consistent and substantial decline across years and venues" when analyzing review quality scores (Q) across ICLR (2018-2025), NeurIPS (2021-2025), and ARR (2022-2025). The analysis reveals that "many of the observed differences are not statistically significant" according to bootstrap testing with 99% confidence intervals and Bonferroni correction.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist empiricism with computational measurement

Author conclusions

The authors conclude that "Contrary to the prevailing narrative, we find no evidence of a decline in review quality. Specifically, we show that the median review quality at computer science conferences has not declined in recent years. Yet, the question demands further investigation, and our follow-up hypotheses outline the problem space for the future study of review quality." They propose five alternative hypotheses (H1-H5) to explain why perceived decline may not reflect actual quality decline, including possibilities that review quality remained stable, that historical data would reveal earlier higher quality, that measurements are incomplete, that worst reviews deteriorated while average remained stable, or that perception of decline is driven by statistical growth in the number of low-quality reviews.

Risk of bias

Selection bias from differential data collection policies: ICLR releases all reviews, NeurIPS includes author opt-in for rejected papers, ARR requires both author and reviewer opt-in with accepted papers only; Survivorship bias toward higher-quality reviews due to opt-in mechanisms at NeurIPS and ARR; Measurement error in LLM-based predictions (ACT κ=0.51, GND κ=0.53) and pragmatic tagging (F1=80.3); Weak correlation (0.21) between proposed metric Q and author-provided quality scores from ACL-2018, suggesting measurement validity concerns; Author bias in ACL-2018 quality scores: authors tend to assign higher scores to reviews giving higher ratings to their papers; Selection bias from data collection policies: NeurIPS and ARR data show imbalance towards accepted papers due to opt-in mechanisms, which may favor higher-quality reviews; Venue comparison bias: Different data availability across venues (ICLR makes all data available, NeurIPS allows opt-in for rejected papers, ARR requires multi-step consent) prevents fair cross-venue comparison; Measurement bias: LLM-based measurements (ACT, GND) have moderate inter-rater agreement (κ² of 0.51 and 0.53) and may not perfectly capture the constructs; Domain and language bias: Study limited to English-language computer science conferences, reducing generalizability; Historical data gap: Earlier review data from before 2021 for NeurIPS and ARR is not publicly available, preventing examination of longer historical trends; Selection bias: Opt-in data collection policies at NeurIPS and ARR favor accepted papers, creating an upper bound on quality estimates; Attrition bias: Different data collection policies across venues prevent fair cross-venue comparison; Measurement error: LLM-based predictions (ACT, GND) have moderate agreement with human scores (κ=0.51-0.53); Confounding: Variation in review forms, guidelines, reviewer expertise, and community culture across venues and time; Missing data: Historical review data prior to 2018 for ICLR and 2021 for NeurIPS not publicly available

Limitations

  • The study acknowledges multiple limitations: "Our study is limited to English as the main language of international academic publishing" and "similarly, our study is limited to computer science conferences in machine learning and AI due to domain familiarity and availability of data." Additionally, "the absence of gold standard data prevents large-scale formal evaluation and iterative improvement" for review itemization
  • Furthermore, "the quality of the predictions that underlie the metrics is likely not perfect, especially for more intricate measurements of review quality like actionability score and grounding score." The authors also note that "differences in data collection policies – but also in paper complexity, time available to review, reviewing load, and differences in community culture – prevent cross-community comparison of review quality."

Open questions raised

  • Lack of gold standard data for review itemization validation and improvement
  • Limited understanding of relationships between reviews and papers, and alignment between review text and scores
  • Need for causal studies on effects of interventions on review quality through controlled experiments
  • Access to historical review data prior to 2021 to test hypothesis that quality was higher earlier
  • Development of additional quality measurement dimensions beyond substantiveness, actionability, and grounding
  • Multi-lingual and cross-domain extensions of review quality measurement
Data: GitHub: https://github.com/UKPLab/arxiv2026-review-quality-estimation; ICLR reviews: OpenReview (CC-BY license); NeurIPS reviews: OpenReview (CC-BY license), opt-in data for rejected papers; ARR reviews: https://arr-data.aclweb.org/ (CC-BY license, multi-step consent); ICLR reviews: accessed via OpenReview (https://openreview.net/); NeurIPS reviews: accessed via OpenReview; ARR (ACL Rolling Review) reviews: https://arr-data.aclweb.org/; Study data and code: https://github.com/UKPLab/arxiv2026-review-quality-estimation; ACL-2018 dataset with author-provided quality scores (Gao et al., 2019); Study data and code; ICLR reviews; NeurIPS reviews; ACL Rolling Review (ARR) dataCode: https://github.com/UKPLab/arxiv2026-review-quality-estimation; GitHub: https://github.com/UKPLab/arxiv2026-review-quality-estimationExtracted from: pdfAgreement 46%

Explore related topics

Related papers