12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Buying the Right to Monitor:Editorial Design in AI-Assisted Peer Review

Zaruhi Hakobyan · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/2
Quality (LMQS)
T
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Three-sided equilibrium model combining game theory, information economics, and organizational design.

Primary method

Analytical derivations under log-concave quality distributions and quadratic polish costs; numerical comparative statics using Gaussian approximation to threshold density and signal-to-noise posterior formulas; implicit function theorem for uniqueness and monotonicity results; truncated bivariate normal distribution theory (Tallis 1961).

Main result

The study identifies three key results: "First, author polishing takes the form of a rat race... Authors therefore spend resources to compete against one another, generating private incentives but social dissipation." Second, "reviewer behavior exhibits a participation transition... the reviewer pool expands, but average evaluative effort falls." Third, "the transition creates a welfare misalignment: Authors benefit from reviewer shirking because the polishing rat race weakens. Editors lose because review signals become less informative." The central managerial implication is that "When AI-assisted reviewing becomes prevalent, journals should not automatically respond by becoming more selective. Tighter selectivity may amplify dissipative author competition without restoring the informativeness of review."

Reports effect sizes.

Research paradigm

Positivist/analytical (economic theory and formal modeling)

Author conclusions

The authors conclude: "The central managerial recommendation is counterintuitive but defensible. When AI-assisted reviewing is rare, conventional editorial intuition is correct: tighter selectivity reduces author rent dissipation. When AI-assisted reviewing is prevalent, conventional intuition reverses: loosening selectivity combined with detection delivers better quality per unit of rent dissipation." They further state: "Generative AI in peer review is best understood not as a threat or a capability but as a technological change that reorganizes the incentive structure of an evaluative organization. The task for organizational design is to update the rules accordingly."

Risk of bias

Model-specification bias: assumptions about functional forms (log-concavity, quadratic polish costs) may not hold empirically; Simplification bias: binary effort choice (0 vs 1) rather than continuous effort spectrum; Assumption bias: specific noise distributions (Gaussian) and cost structures assumed without empirical validation; Detection bias: AI-detection tools exhibit documented false-positive bias against non-native English speakers (acknowledged by authors); Model relies on binary effort choice (read vs. AI-assist) rather than continuous effort, which may oversimplify reviewer behavior; Log-concavity assumption on quality distribution F, while standard, may not hold for all research domains; Symmetric author polishing equilibrium may not capture heterogeneous author responses by career stage or language background; AI-detection modeled as symmetric (uniform false-positive rate) when empirically it exhibits bias against non-native English speakers; Fixed submission mass assumption does not endogenize author entry/exit responses to policy changes; The model assumes symmetric detection without accounting for documented bias against non-native English speakers in AI-detection tools. The treatment of detection as strictly ex-post (not entering reviewer incentives during decision-making) may not reflect institutional reality. The assumption that authors cannot condition polish on privately observed quality (ex-ante equilibrium) simplifies but may miss important strategic behavior.

Limitations

  • The paper acknowledges several limitations: "Allowing polish to condition on privately observed θ would produce a separating equilibrium with richer sorting implications, which we leave to future work." On detection, the authors note "our model treats detection as symmetric: pdet applies uniformly to shirking reports and there is no interaction with author identity
  • In practice, AI-detection tools exhibit documented false-positive bias against non-native English speakers whose writing signatures resemble AI output even when they do not use AI." Additionally, "Full restoration would require instruments outside our policy space: reviewer payments, enforceable AI-use bans, or institutional changes to the reputational system." The treatment of N is restricted: "For tractability and to focus on the new instruments introduced by AI (K and pdet), we hold N fixed at the decentralized post-transition value."

Open questions raised

  • Empirical testing: "The mechanisms we identify—reviewer shirking, author rat-race moderation, and the sign reversal in optimal selectivity—leave observable traces in editorial and bibliometric data. We do not pursue empirical testing here, leaving this as a direction for future work."
  • Separating equilibria: allowing polish to condition on privately observed quality
  • Author heterogeneity: full treatment of author heterogeneity in linguistic profile and asymmetric detection error rates
  • Transfers and payments: allowing monetary transfers to expand feasible frontier
  • Dynamic extensions: endogenizing reviewer reputation across repeated review rounds
  • Multi-journal competition: studying AI-induced tier changes across journal networks
Data: No empirical datasets. The paper uses a numerical baseline calibration (Table 1) but no public datasets are mentioned or made available.Code: Not mentionedExtracted from: pdfAgreement 68%

Explore related topics

Related papers