Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
Vibhhu Sharma, Thorsten Joachims, Sarah Dean · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3805689.3806505
Methodology & findings
Study design
Observational study with secondary analysis of paper-review pairs combined with augmentation using fully LLM-generated reviews.
Sample
N = 125000, 1 group
Primary method
The methodology includes controlling for paper quality through regression analysis and comparative analysis of LLM-generated versus human-authored reviews. Specific statistical software and techniques are not detailed in the abstract.
Main result
The study found that "LLM-assisted reviews seem especially kind to LLM-assisted papers compared to papers with minimal LLM use," but this apparent preference disappears when controlling for paper quality. The analysis reveals that "LLM-assisted reviews are simply more lenient toward lower quality papers in general, and the over-representation of LLM-assisted papers among weaker submissions creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content." Additionally, "fully LLM-generated reviews exhibit severe rating compression that fails to discriminate paper quality, while human reviewers using LLMs substantially reduce this leniency."
Reports effect sizes.
Research paradigm
Empirical-quantitative
Author conclusions
The authors conclude that "these findings provide important input for developing policies that govern the use of LLMs during peer review, and they more generally indicate how LLMs interact with existing decision-making processes." They also note that "meta-reviewers do not merely outsource the decision-making to the LLM" given the differential patterns between human-assisted and fully LLM-generated metareviews.
Risk of bias
Selection bias: LLM-assisted papers may systematically differ in quality from non-LLM papers; Confounding by paper quality: The relationship between LLM use and review leniency is confounded by paper quality; Measurement bias: LLM detection methodology may imperfectly identify LLM use; Temporal bias: Analysis may reflect changing review standards over time; Selection bias in the distribution of LLM-assisted papers across quality tiers; Potential confounding by paper quality when examining interaction effects; Non-random assignment to LLM versus non-LLM reviewers; Observer effects or detection of LLM usage patterns that may not generalize; Selection bias: LLM-assisted papers may be systematically different in quality from non-LLM papers; Confounding by paper quality: The relationship between LLM-assisted reviews and papers is confounded by underlying paper quality; Measurement error: Determining whether papers or reviews are LLM-assisted may have classification errors; Observational design: Causal inference is limited without randomization
Open questions raised
- The authors identify the need for policy development governing LLM use during peer review and suggest that future research should examine how LLMs interact with existing decision-making processes in academic evaluation.
- The paper suggests the need for policy development governing LLM use in peer review and indicates the importance of understanding how LLMs interact with existing decision-making processes in academic publishing.
- The abstract does not explicitly identify future research directions or gaps.
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations