Demanding peer review is associated with higher impact in published science
Huihuang Jiang, Heyang Li, Zifan Wang, Ying Fan, An Zeng · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale computational analysis using fixed-prompt large language model (LLM) pipeline to extract structured reviewer-author interactions from open peer review correspondence.
Sample
N = 8000, 2 groups
Primary method
Spearman rank correlations for continuous metric associations; extreme-group comparisons (top/bottom 5-10% within stratified comparison sets such as year × opinion type); Ordinary Least Squares (OLS) regression with controls for team size, institution count, team average career age, team maximum career age, total review rounds, total reviewers, and year fixed effects; cross-model Pearson correlations for reproducibility validation (Pearson correlations between model pairs reported); keyword-based validation comparing term frequency in metric-specific tails; expert-workflow validation with accuracy measurement. Software: Large language models (Claude Sonnet 4.6, Qwen-3.5-27B, Gemini-3-Flash, DeepSeek-V3.2-nothinking); SciSciNet-V2 for citation impact data (C3 indicator).
Main result
The study found that "stronger criticism, higher-quality comments, and greater revision burden are associated with higher later citation impact within accepted papers." Contrary to intuition that stronger papers pass review more smoothly, papers that attract deeper scrutiny and substantial revision often have greater later impact. Additionally, "review patterns vary little with broad team attributes, consistent with relatively impartial evaluation."
Reports effect sizes.
Research paradigm
Positivist/quantitative empiricism with computational text analysis
Author conclusions
The authors conclude: "Within accepted papers, intense review therefore appears to mark concentrated scrutiny of ambitious or consequential claims, not only manuscript weakness. This changes how difficult review should be interpreted by authors. Receiving strong criticism is not necessarily bad news. In our data, the papers that later become more influential are often those that attract sharper challenges and require more substantial revision." They also note that "peer review is relatively impartial to broad team traits but remains shaped by disciplinary styles of criticism, revision, persuasion, and rebuttal."
Risk of bias
Selection bias: analysis limited to published papers only, excluding rejected manuscripts; Survivor bias: cannot compare accepted vs. rejected papers; Confounding: causality cannot be established; stricter review may not cause higher impact; LLM extraction bias: reliance on fixed-prompt LLM decomposition rather than deterministic parsing; Citation bias: C3 metric may reflect factors other than research quality; Temporal bias: 2022-2024 cohort has shorter observation window for citations; Selection bias: Analysis restricted to accepted papers only, excluding rejected manuscripts that may have experienced different review patterns; Measurement bias: LLM-extracted metrics are inferred from text rather than direct observation of reviewer intent; comment decomposition and response matching are model-inferred rather than deterministic; Outcome measurement bias: C3 (citation count within 3 years) as impact proxy may not capture all forms of scientific influence; Temporal bias: Recent papers (2022-2024) have shorter citation observation windows; Journal-specific bias: Data limited to Nature Communications; findings may not generalize to other journals or disciplines with different review practices; Model dependency: Extracted metrics derived from fixed-prompt LLM procedure; results sensitive to prompt design and model choice; Selection bias: Analysis restricted to published papers only, excluding rejected submissions; Measurement bias: LLM-based extraction may introduce systematic errors in comment decomposition and response matching; Model-dependent bias: Reliance on fixed-prompt LLM outputs rather than deterministic parsing; Citation impact proxy: C3 (3-year citations) may not fully capture long-term impact
Limitations
- "The data include only papers that were ultimately published, so the analysis cannot identify how review differs between accepted and rejected submissions or whether stricter review causes higher impact
- The issue-resolution measure is also an operational proxy based on cross-round persistence rather than a direct measure of reviewer intent." Additionally, the 2022-2024 cohort "can only be observed up to the time covered by that database," limiting citation observation windows for recent papers.
Open questions raised
- The authors identify need for future work including comparisons across journals, rejected-paper trajectories, reviewer-level heterogeneity, and the allocation of scrutiny across different kinds of scientific claims.
- The authors identify the need for: (1) structured account of reviewer-author exchange across rounds at the issue level; (2) comparisons of review between accepted and rejected submissions; (3) causal identification of whether stricter review causes higher impact; (4) process-oriented research on peer review including comparisons across journals; (5) analysis of rejected-paper trajectories; (6) investigation of reviewer-level heterogeneity; (7) examination of allocation of scrutiny across different kinds of scientific claims.
- The authors identify opportunities for future research including: comparisons across journals, rejected-paper trajectories, reviewer-level heterogeneity, and the allocation of scrutiny across different kinds of scientific claims. They also note that "a smaller literature asks whether editorial selection and peer review are related to later citation outcomes, but again at the level of acceptance decisions or manuscript trajectories rather than issue-level interaction."
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations