Buying the Right to Monitor:Editorial Design in AI-Assisted Peer Review
Zaruhi Hakobyan · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Three-sided equilibrium model combining game theory, information economics, and organizational design.
Primary method
Analytical derivations under log-concave quality distributions and quadratic polish costs; numerical comparative statics using Gaussian approximation to threshold density and signal-to-noise posterior formulas; implicit function theorem for uniqueness and monotonicity results; truncated bivariate normal distribution theory (Tallis 1961).
Main result
The study identifies three key results: "First, author polishing takes the form of a rat race... Authors therefore spend resources to compete against one another, generating private incentives but social dissipation." Second, "reviewer behavior exhibits a participation transition... the reviewer pool expands, but average evaluative effort falls." Third, "the transition creates a welfare misalignment: Authors benefit from reviewer shirking because the polishing rat race weakens. Editors lose because review signals become less informative." The central managerial implication is that "When AI-assisted reviewing becomes prevalent, journals should not automatically respond by becoming more selective. Tighter selectivity may amplify dissipative author competition without restoring the informativeness of review."
Reports effect sizes.
Research paradigm
Positivist/analytical (economic theory and formal modeling)
Author conclusions
The authors conclude: "The central managerial recommendation is counterintuitive but defensible. When AI-assisted reviewing is rare, conventional editorial intuition is correct: tighter selectivity reduces author rent dissipation. When AI-assisted reviewing is prevalent, conventional intuition reverses: loosening selectivity combined with detection delivers better quality per unit of rent dissipation." They further state: "Generative AI in peer review is best understood not as a threat or a capability but as a technological change that reorganizes the incentive structure of an evaluative organization. The task for organizational design is to update the rules accordingly."
Risk of bias
Model-specification bias: assumptions about functional forms (log-concavity, quadratic polish costs) may not hold empirically; Simplification bias: binary effort choice (0 vs 1) rather than continuous effort spectrum; Assumption bias: specific noise distributions (Gaussian) and cost structures assumed without empirical validation; Detection bias: AI-detection tools exhibit documented false-positive bias against non-native English speakers (acknowledged by authors); Model relies on binary effort choice (read vs. AI-assist) rather than continuous effort, which may oversimplify reviewer behavior; Log-concavity assumption on quality distribution F, while standard, may not hold for all research domains; Symmetric author polishing equilibrium may not capture heterogeneous author responses by career stage or language background; AI-detection modeled as symmetric (uniform false-positive rate) when empirically it exhibits bias against non-native English speakers; Fixed submission mass assumption does not endogenize author entry/exit responses to policy changes; The model assumes symmetric detection without accounting for documented bias against non-native English speakers in AI-detection tools. The treatment of detection as strictly ex-post (not entering reviewer incentives during decision-making) may not reflect institutional reality. The assumption that authors cannot condition polish on privately observed quality (ex-ante equilibrium) simplifies but may miss important strategic behavior.
Limitations
- The paper acknowledges several limitations: "Allowing polish to condition on privately observed θ would produce a separating equilibrium with richer sorting implications, which we leave to future work." On detection, the authors note "our model treats detection as symmetric: pdet applies uniformly to shirking reports and there is no interaction with author identity
- In practice, AI-detection tools exhibit documented false-positive bias against non-native English speakers whose writing signatures resemble AI output even when they do not use AI." Additionally, "Full restoration would require instruments outside our policy space: reviewer payments, enforceable AI-use bans, or institutional changes to the reputational system." The treatment of N is restricted: "For tractability and to focus on the new instruments introduced by AI (K and pdet), we hold N fixed at the decentralized post-transition value."
Open questions raised
- Empirical testing: "The mechanisms we identify—reviewer shirking, author rat-race moderation, and the sign reversal in optimal selectivity—leave observable traces in editorial and bibliometric data. We do not pursue empirical testing here, leaving this as a direction for future work."
- Separating equilibria: allowing polish to condition on privately observed quality
- Author heterogeneity: full treatment of author heterogeneity in linguistic profile and asymmetric detection error rates
- Transfers and payments: allowing monetary transfers to expand feasible frontier
- Dynamic extensions: endogenizing reviewer reputation across repeated review rounds
- Multi-journal competition: studying AI-induced tier changes across journal networks
Explore related topics
Related papers
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Unlocking the Power of ChatGPT: A Framework for Applying Generative AI in EducationJiahong Su · 2023 · 550 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations