Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time
Ilia Kuznetsov, Rohan Nayak, Alla Rozovskaya, Iryna Gurevych · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational measurement and comparative analysis framework.
Sample
N = 15683, 4 groups
Primary method
Bootstrap hypothesis testing (10,000 iterations, 99% confidence intervals, Bonferroni correction for multiple comparisons per venue). Spearman's rank correlation (ρ) for measuring relationships between measurements. Normalization of count-based and length measurements to [0,1] scale. Sentence-level alignment using all-MiniLM-L6-v2 embeddings with 0.8 similarity threshold. Median comparison across reviewing campaigns as primary descriptive statistic. Human inter-rater agreement assessed via manual annotation and accuracy calculation.
Main result
Contradicting the popular narrative, the study found that "the median review quality at computer science conferences shows no consistent decline" over time. Specifically, the authors report "no consistent and substantial decline across years and venues" when analyzing review quality scores (Q) across ICLR (2018-2025), NeurIPS (2021-2025), and ARR (2022-2025). The analysis reveals that "many of the observed differences are not statistically significant" according to bootstrap testing with 99% confidence intervals and Bonferroni correction.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist empiricism with computational measurement
Author conclusions
The authors conclude that "Contrary to the prevailing narrative, we find no evidence of a decline in review quality. Specifically, we show that the median review quality at computer science conferences has not declined in recent years. Yet, the question demands further investigation, and our follow-up hypotheses outline the problem space for the future study of review quality." They propose five alternative hypotheses (H1-H5) to explain why perceived decline may not reflect actual quality decline, including possibilities that review quality remained stable, that historical data would reveal earlier higher quality, that measurements are incomplete, that worst reviews deteriorated while average remained stable, or that perception of decline is driven by statistical growth in the number of low-quality reviews.
Risk of bias
Selection bias from differential data collection policies: ICLR releases all reviews, NeurIPS includes author opt-in for rejected papers, ARR requires both author and reviewer opt-in with accepted papers only; Survivorship bias toward higher-quality reviews due to opt-in mechanisms at NeurIPS and ARR; Measurement error in LLM-based predictions (ACT κ=0.51, GND κ=0.53) and pragmatic tagging (F1=80.3); Weak correlation (0.21) between proposed metric Q and author-provided quality scores from ACL-2018, suggesting measurement validity concerns; Author bias in ACL-2018 quality scores: authors tend to assign higher scores to reviews giving higher ratings to their papers; Selection bias from data collection policies: NeurIPS and ARR data show imbalance towards accepted papers due to opt-in mechanisms, which may favor higher-quality reviews; Venue comparison bias: Different data availability across venues (ICLR makes all data available, NeurIPS allows opt-in for rejected papers, ARR requires multi-step consent) prevents fair cross-venue comparison; Measurement bias: LLM-based measurements (ACT, GND) have moderate inter-rater agreement (κ² of 0.51 and 0.53) and may not perfectly capture the constructs; Domain and language bias: Study limited to English-language computer science conferences, reducing generalizability; Historical data gap: Earlier review data from before 2021 for NeurIPS and ARR is not publicly available, preventing examination of longer historical trends; Selection bias: Opt-in data collection policies at NeurIPS and ARR favor accepted papers, creating an upper bound on quality estimates; Attrition bias: Different data collection policies across venues prevent fair cross-venue comparison; Measurement error: LLM-based predictions (ACT, GND) have moderate agreement with human scores (κ=0.51-0.53); Confounding: Variation in review forms, guidelines, reviewer expertise, and community culture across venues and time; Missing data: Historical review data prior to 2018 for ICLR and 2021 for NeurIPS not publicly available
Limitations
- The study acknowledges multiple limitations: "Our study is limited to English as the main language of international academic publishing" and "similarly, our study is limited to computer science conferences in machine learning and AI due to domain familiarity and availability of data." Additionally, "the absence of gold standard data prevents large-scale formal evaluation and iterative improvement" for review itemization
- Furthermore, "the quality of the predictions that underlie the metrics is likely not perfect, especially for more intricate measurements of review quality like actionability score and grounding score." The authors also note that "differences in data collection policies – but also in paper complexity, time available to review, reviewing load, and differences in community culture – prevent cross-community comparison of review quality."
Open questions raised
- Lack of gold standard data for review itemization validation and improvement
- Limited understanding of relationships between reviews and papers, and alignment between review text and scores
- Need for causal studies on effects of interventions on review quality through controlled experiments
- Access to historical review data prior to 2021 to test hypothesis that quality was higher earlier
- Development of additional quality measurement dimensions beyond substantiveness, actionability, and grounding
- Multi-lingual and cross-domain extensions of review quality measurement
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations