12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The Shrinking Lifespan of LLMs in Science

Ana Trišović · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale empirical bibliometric analysis drawing on 108,000 papers from Semantic Scholar Academic Graph (S2AG) to track adoption trajectories of 62 language models.

Sample

N = 108000, 3 groups

Primary method

Quadratic regression modeling with heteroskedasticity-robust standard errors (HC1). Delta method for confidence intervals. Three complementary approaches to test inverted-U: (1) quadratic regression with significance testing, (2) joint hypothesis test following Lind & Mehlum (2010), (3) split-sample linear regression following Simonsohn (2018). Log-linear specifications for compression analysis. Cross-sectional regressions with release year fixed effects (Equation 3) for predictors. Inverse-probability weighting using the S2AG citation graph. Bayesian correction combining hand-labeled false positive rates with binomial model of sentence-level classification errors. Implemented in PyFixest.

Main result

The study found that "Scientific adoption of LLMs follows an inverted-U trajectory. Usage of individual models rises after release, peaks, and then declines as newer alternatives enter." Additionally, "The LLM adoption arc is compressing. Successive cohorts of models reach their peak faster and decline sooner, so that the effective window of scientific relevance is shrinking over time. Models released before 2020 accumulated adoption over three to four years; post-2022 models peak within one to two." Quantitatively, "each year of later release is associated with a 27% reduction in time to peak" and "lifespan shrinks by 23% per year of later release."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative bibliometric analysis of observable adoption patterns

Author conclusions

"Taken together, these findings suggest that the pace of language model development is outrunning the timescales of scientific practice. The compression of the model lifecycle has at least three consequences worth highlighting. First, rapid compression raises the practical cost of staying current. When each frontier model remains relevant for only one to two years, researchers face recurring cycles of migration and re-validation. These costs fall unevenly: well-resourced labs can absorb them, while researchers in the social sciences, humanities, or at lower-income institutions face a steeper burden. Second, compression threatens reproducibility, though the severity depends on model openness. For open-weight models, tools such as Ollama, vLLM, and the Hugging Face ecosystem allow researchers to run archived weights indefinitely, decoupling reproducibility from adoption trends. For closed models, no such safeguard exists: providers can deprecate APIs, alter behavior through silent updates, or retire a model on commercial timelines that need not align with scientific ones." The authors further conclude that "bibliometric lifecycle analysis can serve as a demand-side complement to scaling laws."

Risk of bias

Selection bias: Language bias (overrepresentation of English-language publications); Selection bias: Publication type bias (overrepresentation of open-access publications); Measurement bias: Reliance on GPT-4.1-mini for citation classification may introduce systematic misclassification despite Bayesian correction; Statistical power limitations: Small sample size (n=62 models) for subgroup regressions; Measurement bias: Paper-level citations miss informal adoption channels; Temporal bias: Aggregating counts annually smooths within-year dynamics; publication schedules may be uneven; Selection bias: Overrepresentation of English-language and open-access publications in Semantic Scholar and S2ORC; Classification bias: Dependence on zero-shot GPT-4.1-mini for citation sentence classification; potential misclassification noise attenuates estimates conservatively; Sample size limitations: 62 models is relatively small for subgroup regression analyses, limiting statistical power; Measurement bias: Paper-level citations miss informal adoption channels and within-year dynamics are smoothed by annual aggregation; Normalization artifact: Normalizing by each model's peak adoption treats all models as equally important regardless of absolute adoption volume; Classification bias: Reliance on zero-shot GPT-4.1-mini classifier for adoption vs. context distinction; Sample size limitation: Only 62 model variants limits power for subgroup regressions; Coverage bias: Informal adoption channels not captured through paper citations; Temporal bias: Annual aggregation smooths within-year dynamics

Limitations

  • "First, our adoption measure relies on Semantic Scholar and S2ORC, which overrepresent English-language, open-access publications
  • Our inverse-probability weighting partially corrects for this, but adoption patterns in fields with lower open-access rates may differ
  • Second, our classification of citation sentences into context versus adoption depends on a zero-shot GPT-4.1-mini classifier
  • Although we validate against hand-labeled data and apply a Bayesian false-positive correction, misclassification noise likely attenuates our estimates, making the patterns we report conservative rather than inflated
  • Third, with 62 model variants our sample is small for subgroup regressions, and some null results may reflect limited power rather than genuine absence of effects
  • Fourth, we measure scientific adoption through paper-level citations, which capture formal scholarly outputs but miss informal adoption channels

Open questions raised

  • Despite growing awareness of LLM adoption in research, surprisingly little is known about what happens after adoption
  • The field lacks understanding of adoption dynamics across model generations and properties predicting scientific utility
  • No prior study has characterized the full adoption arc of a specific model or measured whether adoption arcs are compressing
  • Informal adoption channels are not captured in bibliometric measures
  • The research raises a fundamental question: is the rapid turnover of LLMs compatible with accumulation of reliable, replicable scientific knowledge?
  • The authors identify the following gaps: (1) lack of understanding about what happens after LLM adoption in scientific workflows; (2) absence of characterization of adoption patterns - whether usage accumulates, plateaus, or declines - and how this varies across model generations; (3) limited knowledge about which model properties predict longevity versus rapid displacement; (4) insufficient understanding of whether rapid LLM turnover is compatible with accumulation of reliable, replicable scientific knowledge; and (5) need for new norms around model versioning, archiving, and citation practices.
Data: Semantic Scholar Academic Graph (S2AG) - February 2026 snapshot; S2ORC corpus (S2 Open Research Corpus) for full-text extraction; S2ORC corpus for full-text extractionCode: Not mentionedExtracted from: pdfAgreement 60%

Explore related topics

Related papers