12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review

Vibhhu Sharma, Thorsten Joachims, Sarah Dean · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3805689.3806505

Methodology & findings

Study design

Observational study with secondary analysis of paper-review pairs combined with augmentation using fully LLM-generated reviews.

Sample

N = 125000, 1 group

Primary method

The methodology includes controlling for paper quality through regression analysis and comparative analysis of LLM-generated versus human-authored reviews. Specific statistical software and techniques are not detailed in the abstract.

Main result

The study found that "LLM-assisted reviews seem especially kind to LLM-assisted papers compared to papers with minimal LLM use," but this apparent preference disappears when controlling for paper quality. The analysis reveals that "LLM-assisted reviews are simply more lenient toward lower quality papers in general, and the over-representation of LLM-assisted papers among weaker submissions creates a spurious interaction effect rather than genuine preferential treatment of LLM-generated content." Additionally, "fully LLM-generated reviews exhibit severe rating compression that fails to discriminate paper quality, while human reviewers using LLMs substantially reduce this leniency."

Reports effect sizes.

Research paradigm

Empirical-quantitative

Author conclusions

The authors conclude that "these findings provide important input for developing policies that govern the use of LLMs during peer review, and they more generally indicate how LLMs interact with existing decision-making processes." They also note that "meta-reviewers do not merely outsource the decision-making to the LLM" given the differential patterns between human-assisted and fully LLM-generated metareviews.

Risk of bias

Selection bias: LLM-assisted papers may systematically differ in quality from non-LLM papers; Confounding by paper quality: The relationship between LLM use and review leniency is confounded by paper quality; Measurement bias: LLM detection methodology may imperfectly identify LLM use; Temporal bias: Analysis may reflect changing review standards over time; Selection bias in the distribution of LLM-assisted papers across quality tiers; Potential confounding by paper quality when examining interaction effects; Non-random assignment to LLM versus non-LLM reviewers; Observer effects or detection of LLM usage patterns that may not generalize; Selection bias: LLM-assisted papers may be systematically different in quality from non-LLM papers; Confounding by paper quality: The relationship between LLM-assisted reviews and papers is confounded by underlying paper quality; Measurement error: Determining whether papers or reviews are LLM-assisted may have classification errors; Observational design: Causal inference is limited without randomization

Open questions raised

  • The authors identify the need for policy development governing LLM use during peer review and suggest that future research should examine how LLMs interact with existing decision-making processes in academic evaluation.
  • The paper suggests the need for policy development governing LLM use in peer review and indicates the importance of understanding how LLMs interact with existing decision-making processes in academic publishing.
  • The abstract does not explicitly identify future research directions or gaps.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 56%

Explore related topics

Related papers