Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI
Wenqing Wu, Chengzhi Zhang, Yi Zhao, Tong Bao · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Observational study with longitudinal document analysis.
Sample
N = 188198, 4 groups
Primary method
Descriptive statistics (mean, distribution analysis); Lexical complexity measurement via TAALES tool (400+ indices); Syntactic complexity analysis via TAASSC tool (5 selected metrics: advcl per cl, nsubj per cl, mark per cl, aux per cl, dobj per cl); Aspect identification via pre-trained sequence labeling model with BERT embeddings and multilayer perceptron classification (Equation 1: softmax activation); Maximum likelihood estimation (MLE) for LLM-assisted text detection (Equation 2); Spearman rank correlation coefficient analysis between aspect mentions and reviewer scores/confidence scores; Scatter plot visualization with simple linear regression fits (noted as 'visual aid for qualitative comparison rather than formal regression analysis')
Main result
The results indicate that "following the emergence of LLMs, peer review texts have become longer and more fluent, with increased emphasis on summaries and surface-level clarity, as well as more standardized linguistic patterns, particularly reviewers with lower confidence score." Additionally, "attention to deeper evaluative dimensions, such as originality, replicability, and nuanced critical reasoning, has declined."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude: "Overall, our findings indicate that the emergence of LLMs is associated with longer and more fluent peer-review texts, increased emphasis on summaries and surface-level clarity, and increasingly standardized linguistic patterns effects that are particularly pronounced among reviewers with lower self-reported confidence. At the same time, attention to deeper evaluative dimensions, including originality, replicability, and fine-grained critical reasoning, has declined." They further state: "Although LLM-assisted reports show a modest improvement in the informativeness of recommendations, this trend raises concerns about a potential trade-off between linguistic fluency and evaluative depth in peer review."
Risk of bias
Selection bias: Data limited to two top-tier AI conferences (ICLR, NeurIPS); findings may not generalize to other fields or conferences; Composition bias: NeurIPS data predominantly comprises accepted papers while ICLR has more rejected papers; this compositional difference may affect results on length, aspect mentions, and LLM assistance estimates; Detection uncertainty: LLM-assisted detection method identifies 'may have' LLM assistance rather than definitive identification; uses lexicon-based filtering which may miss LLM-assisted reviews or generate false positives; Temporal confounders: Changes in review length/patterns could be driven by factors other than LLM adoption (e.g., increasing paper complexity, reviewer pool changes); Aspect annotation bias: Relies on pre-trained model with 92.75% accuracy; systematic errors in aspect identification would propagate through analyses; Observational design: Cannot establish causation; associations may reflect reverse causation or unmeasured confounders; Selection bias: Data limited to top AI conferences (ICLR, NeurIPS); limited generalizability to other fields; Imbalanced dataset: NeurIPS data primarily comprises accepted papers while ICLR contains more rejected papers; this difference may influence results such as average length and aspect mentions; Detection uncertainty: LLM-assisted detection method identifies reports that 'may have' LLM assistance but does not provide absolute certainty; Observational bias: Study based on correlational analysis, not causal inference; cannot establish causation; Missing data: NeurIPS 2020 data not available from OpenReview; Selection bias: Data primarily from top AI conferences (ICLR, NeurIPS), not representative of all fields; Data composition bias: NeurIPS reviews predominantly from accepted papers; ICLR contains more rejected papers; Detection uncertainty: LLM-assisted detection method does not provide absolute certainty; Confounding: Cannot establish causality due to observational design; multiple factors may drive changes in review text beyond LLM usage; Temporal bias: Missing NeurIPS 2020 data; Domain specificity: Results may not generalize to other academic fields
Limitations
- The authors state that "our data is primarily drawn from top conferences in the fields of machine learning and deep learning
- To generalize these findings to other fields, corresponding data from those areas would be required
- The NeurIPS data we used primarily comprises accepted papers, while the ICLR data contains more rejected papers than accepted ones
- This difference may influence certain results, such as average length, aspect mentions, and potential LLM assistance usage." Additionally, "the LLM-assisted detection method we used identifies reports that may have LLM assistance but does not provide absolute certainty" and "this study focuses on identifying correlations and observed patterns rather than establishing causation
- Although our analyses highlight associations between review text, review aspect, sentiment, overall scores, confidence scores, and LLM-assisted, they do not imply causal relationships due to the observational nature of our study."
Open questions raised
- Causal mechanisms: Future work should employ experimental or quasi-experimental methodologies to explore causal relationships between LLM usage and changes in peer review comments
- Interdisciplinary generalization: Analyses should broaden scope to peer review reports from leading journals such as PLOS ONE and Nature Communications to identify trends across different fields and publication standards
- Advanced detection methods: More sophisticated tools and techniques are needed for detecting LLM usage in peer review comments, including advanced text analysis methods and LLM-based detection systems
- Review quality assessment: Whether LLM assistance leads to more accurate, insightful, or useful reviews remains an open question requiring direct measurement combining linguistic analyses with expert evaluations of review helpfulness, accuracy, and constructiveness
- Deeper syntactic analysis: The study did not examine detailed word usage or conduct more in-depth syntactic analyses such as changes in syntactic dependencies
- Whether LLMs are altering the core evaluative functions of peer review remains unclear
Explore related topics
Related papers
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Publishers’ and journals’ instructions to authors on use of generative artificial intelligence in academic and scientific publishing: bibliometric analysisConner Ganjavi · 2024 · 199 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations