12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How much are LLMs changing the language of academic papers after ChatGPT? A multi-database and full text analysis

Scientometrics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
1
Citations
7.11
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05601-5

Methodology & findings

Study design

Multi-database bibliometric analysis combined with full-text computational analysis.

Sample

N = 2400000, 15 groups

Primary method

Proportion normalization: Term counts divided by total publications per year to account for database growth; Geometric mean calculation: Used for average usage of LLM-associated terms (instead of arithmetic mean due to highly skewed data with many zero occurrences); Chi-square tests (χ² test, p < 0.01): Used to screen candidate terms for co-occurrence with seed terms; Pearson correlation analysis: Applied to measure correlations between LLM-associated term frequencies in PMC full texts; Conditional probability calculations: For assessing co-occurrence of term pairs within same papers; Error bar analysis: Overlapping vs. non-overlapping error bars used to assess statistical significance at 95% confidence level

Main result

The study found that "underscore[s/d/ing] was the most frequent term, occurring in about 20% of PMC open-access publications, rising from just 3.6% in 2022, which represents almost a 5.5 × increase in only two years." Additionally, "delve increased by up to 1,582% in Scopus and 1,331% in WoS, while underscore increased by 1,046% and 895%, respectively" between 2022 and 2024. The analysis also revealed that "in 2024, 59.3% of PMC papers mentioning delve (37,534 in total) also mentioned underscore (22,241 papers)," compared with just 1.3% in 2022, indicating substantially increased co-occurrence of LLM-associated terms after ChatGPT's release.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/Empiricist - quantitative measurement of linguistic patterns in large databases

Author conclusions

"In answer to RQ1, the use of LLM-associated terms has increased sharply across scholarly databases after ChatGPT's public release in late 2022. Delve (+ 1,360%) and underscore (+ 1,062%) grew much faster than traditional terms such as investigate (+ 8.9%) and highlight (+ 85%) between 2022 and 2024. By 2024, underscore appeared in about 20% of PMC full-text papers and 11% of Dimensions papers, while pivotal reached 15% in PMC. There are also clear disciplinary differences. The growth of LLM terms was highest in STEM fields, often exceeding 3,000%. In contrast, there were smaller increases in the social sciences and arts and humanities, mostly below 500%." The authors further note that "the increasing prevalence of AI-related language in full texts as well as titles and abstracts may be a welcome development overall. It suggests that LLMs are reducing the language barrier to academic publishing for non-English speakers and hence partly addressing the current unacceptable level of global inequality in science publishing."

Risk of bias

Selection bias: Limited to English-language terms; non-English articles with translated abstracts could skew results; Non-representative sample: PMC analysis limited to open-access biomedical/life-science publications only; excludes major paywalled journals (NEJM, Lancet, JAMA); Confounding: Some terms may be discipline-specific (e.g., 'foster' in education, 'heighten' in psychology) independent of LLM influence; Temporal confounding: Cannot distinguish LLM influence from other long-term trends or changes in editorial practices; Language bias: Analysis focuses on English-language academic writing; non-native English speakers may use LLMs differently for translation/proofreading; Database variation: OpenAlex uses first publication date (often preprints) rather than formal publication date, appearing about 1 year ahead of other databases; Selection bias: Analysis limited to six databases; may overlook other academic databases or non-indexed publications; Language bias: Analysis restricted to English-language terms; non-English papers with translated English abstracts may introduce artifacts; Temporal bias: Normalization by publication volume helps mitigate, but absolute growth in publication output may confound results; Disciplinary bias: Different fields have varying rates of LLM adoption; STEM fields show much higher uptake than social sciences/humanities; Open-access bias: PMC analysis covers only open-access biomedical/life science publications, excluding paywalled journals like The Lancet, JAMA; Term selection bias: 12 terms chosen based on prior literature and exploratory analysis; other LLM-associated terms may exist but were not included; Causality assumption: Study cannot definitively prove LLMs cause increased term usage; could reflect editorial changes, author preferences, or other factors; Selection bias: Only open-access PMC papers analyzed, excluding many leading medical journals (The New England Journal of Medicine, The Lancet, JAMA); Language bias: Analysis limited to English-language terms and abstracts; non-English articles with translated English abstracts could artificially inflate LLM-term counts; Discipline-specific bias: Some terms naturally occur in certain disciplines (e.g., 'foster' in education, 'heighten' in psychology) independent of LLM use; Temporal confounding: Study cannot distinguish whether increases reflect direct LLM text generation, editorial/publishing practice changes linked to LLM use, or longer-term trends unrelated to LLMs; Database heterogeneity: Different databases use different indexing practices; OpenAlex uses first publication date (often preprint) rather than formal publication date

Limitations

  • "The results are limited by the set of 12 LLM-associated terms used and the six databases analysed, and we may have overlooked other terms commonly suggested by LLMs, including those used by models other than ChatGPT." Additionally, "some of the studied terms may naturally be used more frequently in specific disciplines, independent of LLM usage
  • For example, the term foster or fostering may commonly appear in fields like education, social work, child development or animal care, making these counts less reliable as direct indicators of LLM influence." Furthermore, "The PMC full-text dataset is available only for open-access biomedical and life-science research, and not all articles are included in the analysis
  • As a result, the findings mainly reflect LLM-term usage in open-access publications and may not fully represent non-open-access journals." The authors also note that "this study analysed changes in the frequency and co-occurrence of selected LLM-associated words in academic publications but did not assess potential semantic or stylistic changes in sentence structure which may also be influenced by LLMs."

Open questions raised

  • How different LLMs (ChatGPT, Gemini, DeepSeek) differ in their influence on academic writing
  • Whether patterns reflect direct LLM text generation or indirect changes in editorial/publishing practices
  • Language patterns in non-English academic publications
  • Semantic and stylistic changes in sentence structure beyond word frequency
  • Analysis of a broader range of alternative terms beyond those tested
  • Investigation of whether LLM influence differs across specific subject disciplines
Data: PMC full-text frequency dataset (2021-July 2025); PMC full-text term frequency dataset (2021-July 2025)Code: Webometric Analyst software; GitHub: https://github.com/MikeThelwall/Webometric_AnalystExtracted from: pdfAgreement 51%

Explore related topics

Related papers