12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI-assisted writing and the reorganization of scientific knowledge

Erjia Yan, Chaoqun Ni · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Observational panel study using two million full-text research articles published 2021-2024 linked to citation networks.

Sample

N = 2000000, 6 groups

Primary method

Fixed-effects panel models with author×field and year fixed effects; Frisch-Waugh-Lovell (FWL) residualization for two-way fixed effects estimation; Year-specific interaction models (allowing relationship to vary flexibly by year); First-difference models (within-author acceleration models); Heteroskedasticity-robust (HC1) standard errors; Clustered standard errors (author-field level) for robustness checks; Interaction models with continuous academic age and field-level entropy; Placebo tests using pre-period data (2021-2022); Software: Python (Statsmodels OLS), BigQuery for citation computations

Main result

The study found that "after the diffusion of large language models in 2023, AI-assisted writing intensity became positively associated with citation disruption." The authors document that "in 2023 and 2024, it becomes positive and remains so," indicating that "papers with higher AI-assisted writing are increasingly associated with more disruptive citation structures." Simultaneously, "the post-2023 increase in disruption is not accompanied by a corresponding expansion in cross-field recombination," suggesting that "generative AI is associated with more disruptive citation structures without a corresponding expansion in cross-field recombination."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative observational study using citation networks as observable indicators of scientific knowledge organization

Author conclusions

The authors conclude: "If these patterns persist, generative AI may alter the balance between consolidation and displacement in science by reshaping how prior knowledge is recombined rather than simply by accelerating writing." They also note: "generative AI is associated with a reconfiguration of citation structure: papers with higher AI-writing intensity become more disruptive, but not because they draw on a broader set of fields. Instead, the evidence is more consistent with changing patterns of recombination within relatively narrower knowledge inputs."

Risk of bias

Observational design prevents causal inference; unobserved confounders (journal policies, editorial norms, topic changes) may drive results; Sample restricted to PubMed Central articles, enriched in biomedical research, may not represent all scientific output; AI-assisted writing measure is classifier-based probabilistic proxy, not direct tool usage; measurement error possible; Assignment of AI exposure to all coauthors (not just writers) may attenuate true effects; Selection bias from authors choosing whether to use AI tools; Potential stylistic confounding—changes in academic writing norms unrelated to AI use; Selection bias from PMC sample restriction (enriched in biomedical research, may not be fully representative of all scientific output); Measurement error: AI-assisted writing intensity is a probabilistic proxy based on classifier, not direct tool usage measurement; Confounding from unobserved changes in journals, editorial norms, or research topics; Attenuation bias from assigning AI exposure to all coauthors regardless of their role in manuscript preparation; Potential stylistic confounding: authors acknowledge residual stylistic confounding and under-detection of heavily revised or non-textual uses cannot be fully excluded; Temporal confounding: diffusion of generative AI coincides with other potential changes in scientific practices; Selection bias: Sample restricted to PubMed Central articles, which may not be fully representative of all scientific output; Measurement error: AI-assisted writing intensity is a classifier-based proxy, not direct tool usage recording; Confounding: Unobserved changes in journal practices, editorial norms, or research topics; Attrition/truncation bias: Addressed through fixed forward-citation windows to ensure equal exposure; Composition bias: Mitigated through within-author, within-field identification strategy

Limitations

  • The authors state: "First, the analysis is observational and cannot establish a causal effect of AI-assisted writing on scientific outcomes
  • Unobserved changes in journals, fields, editorial norms, or research topics may also contribute to the observed pattern
  • Second, our measure of AI-assisted writing intensity is a classifier-based proxy and should not be interpreted as a direct record of tool use
  • Third, the disruption index captures structural relationships in citation networks rather than the substantive content or long-run importance of ideas."

Open questions raised

  • Whether observed citation disruption reflects genuine scientific creativity or merely structural reorganization without substantive value
  • Long-run importance and durability of disruptive papers produced with AI assistance
  • Substantive content analysis of AI-assisted papers versus citation network structure alone
  • Mechanisms linking AI use to specific patterns of knowledge recombination within versus across fields
  • Causal effects of AI-assisted writing (vs. observational associations)
  • The paper identifies that despite increased disruption post-2023, this is not accompanied by broader cross-field sourcing, suggesting a need to understand mechanisms of recombination within narrower knowledge domains
Data: OpenAlex Walden snapshot with PubMed Central full-text articles; OpenAlex (Walden snapshot, updated March 6, 2026) - https://doi.org/10.5281/zenodo.19581823 (Zenodo repository); Zenodo datasetCode: GitHub: https://github.com/erjiayan/AIdisruptiveness; GitHub: https://github.com/erjiayan/AIdisruptiveness; GitHubExtracted from: pdfAgreement 52%

Explore related topics

Related papers