12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI

Wenqing Wu, Chengzhi Zhang, Yi Zhao, Tong Bao · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Observational study with longitudinal document analysis.

Sample

N = 188198, 4 groups

Primary method

Descriptive statistics (mean, distribution analysis); Lexical complexity measurement via TAALES tool (400+ indices); Syntactic complexity analysis via TAASSC tool (5 selected metrics: advcl per cl, nsubj per cl, mark per cl, aux per cl, dobj per cl); Aspect identification via pre-trained sequence labeling model with BERT embeddings and multilayer perceptron classification (Equation 1: softmax activation); Maximum likelihood estimation (MLE) for LLM-assisted text detection (Equation 2); Spearman rank correlation coefficient analysis between aspect mentions and reviewer scores/confidence scores; Scatter plot visualization with simple linear regression fits (noted as 'visual aid for qualitative comparison rather than formal regression analysis')

Main result

The results indicate that "following the emergence of LLMs, peer review texts have become longer and more fluent, with increased emphasis on summaries and surface-level clarity, as well as more standardized linguistic patterns, particularly reviewers with lower confidence score." Additionally, "attention to deeper evaluative dimensions, such as originality, replicability, and nuanced critical reasoning, has declined."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude: "Overall, our findings indicate that the emergence of LLMs is associated with longer and more fluent peer-review texts, increased emphasis on summaries and surface-level clarity, and increasingly standardized linguistic patterns effects that are particularly pronounced among reviewers with lower self-reported confidence. At the same time, attention to deeper evaluative dimensions, including originality, replicability, and fine-grained critical reasoning, has declined." They further state: "Although LLM-assisted reports show a modest improvement in the informativeness of recommendations, this trend raises concerns about a potential trade-off between linguistic fluency and evaluative depth in peer review."

Risk of bias

Selection bias: Data limited to two top-tier AI conferences (ICLR, NeurIPS); findings may not generalize to other fields or conferences; Composition bias: NeurIPS data predominantly comprises accepted papers while ICLR has more rejected papers; this compositional difference may affect results on length, aspect mentions, and LLM assistance estimates; Detection uncertainty: LLM-assisted detection method identifies 'may have' LLM assistance rather than definitive identification; uses lexicon-based filtering which may miss LLM-assisted reviews or generate false positives; Temporal confounders: Changes in review length/patterns could be driven by factors other than LLM adoption (e.g., increasing paper complexity, reviewer pool changes); Aspect annotation bias: Relies on pre-trained model with 92.75% accuracy; systematic errors in aspect identification would propagate through analyses; Observational design: Cannot establish causation; associations may reflect reverse causation or unmeasured confounders; Selection bias: Data limited to top AI conferences (ICLR, NeurIPS); limited generalizability to other fields; Imbalanced dataset: NeurIPS data primarily comprises accepted papers while ICLR contains more rejected papers; this difference may influence results such as average length and aspect mentions; Detection uncertainty: LLM-assisted detection method identifies reports that 'may have' LLM assistance but does not provide absolute certainty; Observational bias: Study based on correlational analysis, not causal inference; cannot establish causation; Missing data: NeurIPS 2020 data not available from OpenReview; Selection bias: Data primarily from top AI conferences (ICLR, NeurIPS), not representative of all fields; Data composition bias: NeurIPS reviews predominantly from accepted papers; ICLR contains more rejected papers; Detection uncertainty: LLM-assisted detection method does not provide absolute certainty; Confounding: Cannot establish causality due to observational design; multiple factors may drive changes in review text beyond LLM usage; Temporal bias: Missing NeurIPS 2020 data; Domain specificity: Results may not generalize to other academic fields

Limitations

  • The authors state that "our data is primarily drawn from top conferences in the fields of machine learning and deep learning
  • To generalize these findings to other fields, corresponding data from those areas would be required
  • The NeurIPS data we used primarily comprises accepted papers, while the ICLR data contains more rejected papers than accepted ones
  • This difference may influence certain results, such as average length, aspect mentions, and potential LLM assistance usage." Additionally, "the LLM-assisted detection method we used identifies reports that may have LLM assistance but does not provide absolute certainty" and "this study focuses on identifying correlations and observed patterns rather than establishing causation
  • Although our analyses highlight associations between review text, review aspect, sentiment, overall scores, confidence scores, and LLM-assisted, they do not imply causal relationships due to the observational nature of our study."

Open questions raised

  • Causal mechanisms: Future work should employ experimental or quasi-experimental methodologies to explore causal relationships between LLM usage and changes in peer review comments
  • Interdisciplinary generalization: Analyses should broaden scope to peer review reports from leading journals such as PLOS ONE and Nature Communications to identify trends across different fields and publication standards
  • Advanced detection methods: More sophisticated tools and techniques are needed for detecting LLM usage in peer review comments, including advanced text analysis methods and LLM-based detection systems
  • Review quality assessment: Whether LLM assistance leads to more accurate, insightful, or useful reviews remains an open question requiring direct measurement combining linguistic analyses with expert evaluations of review helpfulness, accuracy, and constructiveness
  • Deeper syntactic analysis: The study did not examine detailed word usage or conduct more in-depth syntactic analyses such as changes in syntactic dependencies
  • Whether LLMs are altering the core evaluative functions of peer review remains unclear
Data: ICLR and NeurIPS peer review reports; Code and dataset for this paper; ICLR and NeurIPS peer review data; ICLR reviews; NeurIPS reviewsCode: GitHub: https://github.com/njust-winchy/LLM_impact; GitHub: https://github.com/njust-winchy/LLM_impact; GitHub: https://github.com/njust-winchy/LLM_impactExtracted from: pdfAgreement 56%

Explore related topics

Related papers