12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Publish and Perish: How AI-Accelerated Writing Without Proportional Verification Investment Degrades Scientific Knowledge

Seok Joon Kwon · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Minimal dynamical systems model with coupled differential equations.

Main result

The study found that "knowledge output crosses below K0 at t = 6.1 yr (year 2028), marking paradox onset" and "by t = 20 yr (year 2042), K/K0 = 0.68", demonstrating a 32% knowledge loss. The model predicts "a two-phase trajectory" where "knowledge output initially increases, peaking at 1.10K0 at t = 3.5 yr (circa year 2026)" during a deceptive honeymoon period, but subsequently "K declines monotonically as quality erosion overwhelms throughput gains." Empirical validation shows "NeurIPS main track grew from 9,467 to 21,575 submissions (+128%, 2020-2025)" and "ICLR grew from 2,594 to 19,631 (+657%, 2020-2026)", with "NeurIPS grew at approximately 5% compound annual growth rate (CAGR) during 2020-2022 (pre-ChatGPT), accelerating to approximately 27% CAGR during 2022-2025 (post-ChatGPT)".

Research paradigm

Systems theory and mathematical modeling of knowledge production systems

Author conclusions

The authors conclude: "The counterintuitive lesson is clear: in science, faster writing without proportional verification investment creates less knowledge, not more. The choice facing the research community is whether to invest deliberately in the bottleneck now, or to accept a future where we publish prolifically while knowing less." They also state that "The paradox is not inevitable. Escaping it requires investing in the bottleneck: review infrastructure that amplifies human judgment (increasing δ), institutional quality standards that raise the minimum quality floor qmin, and metric reform that values verification depth over publication volume." Finally, they emphasize: "Neither technology investment nor regulation alone suffices; only the combination breaks the feedback loop linking queue pressure to quality erosion."

Risk of bias

Confounding factors in submission growth (community expansion, industry participation, conference scope broadening); Homogeneous population assumption masks discipline-specific and career-stage variation; AI text detector limitations (known false positive and negative rates) in validating reviewer AI adoption; Selection bias in empirical data (AI-intensive venues vs. biology venues show differential patterns); Model calibration may overstate AI writing tool contribution (γ = 2.0 likely overstates compared to Amdahl's law lower bound of 0.3-0.7); Model calibration uncertainty: γ = 2.0 may overstate AI-tool-specific contribution to submission growth, conflating AI effects with community expansion and scope broadening; Homogeneity assumption: Model does not account for discipline-specific, career-stage-specific, or institutional variations in AI adoption; Confounding factors in empirical validation: Post-ChatGPT acceleration at AI/ML venues reflects multiple trends beyond writing tools (industry investment, community expansion); Quality measurement: Aggregated scalar q may mask differential degradation rates across review dimensions; Unvalidated core prediction: Quality trajectory q(t) lacks direct empirical validation despite being central to the model; Selection bias in venue comparison: bioRxiv deceleration may reflect user base saturation rather than lower AI adoption; Attribution bias: Submission growth at AI/ML venues reflects multiple concurrent factors (community expansion, industry investment, conference scope broadening) beyond AI writing tools; model may overstate AI tools' contribution (γ = 2.0); Confounding: Differential acceleration pattern (AI-intensive venues accelerating post-ChatGPT while biology venues decelerate) provides suggestive but not definitive causal evidence; definitive test would require within-venue variation in AI tool access not currently available; Model validation gap: Central prediction (verification quality trajectory q(t)) lacks direct empirical validation; only submission data and reviewer AI adoption rates are validated; Parameter estimation uncertainty: Model has 5 parameters explored via sensitivity analysis but with limited empirical constraints; critical parameters (Qc, λ, μ, η, qmin) derived from plausibility arguments rather than direct measurement; Homogeneity assumption: Model assumes homogeneous researcher and reviewer populations, not capturing discipline-specific, career-stage, or geographic/institutional vulnerabilities; Temporal generalization: Empirical data spans 2008-2026; predictions extend to 2042 (20-year horizon) with unknown validity; Desktop review data limitation: AI-generated review estimates rely on GPT-text detectors with known false positive/negative rates; true prevalence uncertain

Limitations

  • The authors state: "the model assumes homogeneous researcher and reviewer populations" and note that "the scalar value q aggregates multi-aspect review (methodological rigor, novelty assessment, statistical validity, ethical review) that may degrade at different rates under AI pressure." Additionally, "the model assumes passive parameter evolution, not capturing strategic gaming dynamics" and "submission growth at AI/ML venues reflects multiple concurrent trends beyond AI writing tools: the AI research community has expanded enormously since 2020, industry R&D investment has surged, and conferences have broadened their scope." Furthermore, "the verification quality trajectory q(t) is the model's central prediction but lacks direct empirical validation: while submission growth S(t) is validated against venue data and AI adoption varies by discipline (physics versus biology), career stage (junior versus senior), reviewer AI adoption φr is validated against Liang et al.[1,2], no comparable time-series data exists for aggregate review depth."

Open questions raised

  • Direct empirical validation of verification quality trajectory q(t) over time
  • Time-series data on aggregate review depth, retraction rates, post-publication correction frequencies, review length and substantiveness
  • Reproducibility metrics tracked longitudinally over time
  • Disaggregated models stratified by field to reveal discipline-specific vulnerabilities
  • Multi-aspect review degradation at different rates under AI pressure
  • Strategic gaming dynamics (authors overloading competitors' queues, weaponizing AI reviews, publishers exploiting AI for volume)
Data: NeurIPS submission data (2008-2025) from [19, 20]; ICLR submission data (2013-2026) from [21]; arXiv monthly submissions data (2008-2025) from [22]; bioRxiv annual preprints data (2014-2025) from [23]; ICLR AI-generated review data (2024-2025) from [1, 2]; NeurIPS submission data (2008-2025) - cited as [19, 20]; ICLR submission data (2013-2026) - cited as [21]; arXiv monthly submissions (2008-2025) - cited as [22]; bioRxiv annual preprints (2014-2025) - cited as [23]; NeurIPS submission data (2008-2025): citations [19, 20]; ICLR submission data (2013-2026): citation [21]; arXiv monthly submission data (2008-2025): citation [22]; bioRxiv annual preprint data (2014-2025): citation [23]; ICLR AI-generated review data (2024-2025): citations [1, 2]Extracted from: pdfAgreement 58%

Explore related topics

Related papers