How much are LLMs changing the language of academic papers after ChatGPT? A multi-database and full text analysis
Scientometrics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05601-5
Methodology & findings
Study design
Multi-database bibliometric analysis combined with full-text computational analysis.
Sample
N = 2400000, 15 groups
Primary method
Proportion normalization: Term counts divided by total publications per year to account for database growth; Geometric mean calculation: Used for average usage of LLM-associated terms (instead of arithmetic mean due to highly skewed data with many zero occurrences); Chi-square tests (χ² test, p < 0.01): Used to screen candidate terms for co-occurrence with seed terms; Pearson correlation analysis: Applied to measure correlations between LLM-associated term frequencies in PMC full texts; Conditional probability calculations: For assessing co-occurrence of term pairs within same papers; Error bar analysis: Overlapping vs. non-overlapping error bars used to assess statistical significance at 95% confidence level
Main result
The study found that "underscore[s/d/ing] was the most frequent term, occurring in about 20% of PMC open-access publications, rising from just 3.6% in 2022, which represents almost a 5.5 × increase in only two years." Additionally, "delve increased by up to 1,582% in Scopus and 1,331% in WoS, while underscore increased by 1,046% and 895%, respectively" between 2022 and 2024. The analysis also revealed that "in 2024, 59.3% of PMC papers mentioning delve (37,534 in total) also mentioned underscore (22,241 papers)," compared with just 1.3% in 2022, indicating substantially increased co-occurrence of LLM-associated terms after ChatGPT's release.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Empiricist - quantitative measurement of linguistic patterns in large databases
Author conclusions
"In answer to RQ1, the use of LLM-associated terms has increased sharply across scholarly databases after ChatGPT's public release in late 2022. Delve (+ 1,360%) and underscore (+ 1,062%) grew much faster than traditional terms such as investigate (+ 8.9%) and highlight (+ 85%) between 2022 and 2024. By 2024, underscore appeared in about 20% of PMC full-text papers and 11% of Dimensions papers, while pivotal reached 15% in PMC. There are also clear disciplinary differences. The growth of LLM terms was highest in STEM fields, often exceeding 3,000%. In contrast, there were smaller increases in the social sciences and arts and humanities, mostly below 500%." The authors further note that "the increasing prevalence of AI-related language in full texts as well as titles and abstracts may be a welcome development overall. It suggests that LLMs are reducing the language barrier to academic publishing for non-English speakers and hence partly addressing the current unacceptable level of global inequality in science publishing."
Risk of bias
Selection bias: Limited to English-language terms; non-English articles with translated abstracts could skew results; Non-representative sample: PMC analysis limited to open-access biomedical/life-science publications only; excludes major paywalled journals (NEJM, Lancet, JAMA); Confounding: Some terms may be discipline-specific (e.g., 'foster' in education, 'heighten' in psychology) independent of LLM influence; Temporal confounding: Cannot distinguish LLM influence from other long-term trends or changes in editorial practices; Language bias: Analysis focuses on English-language academic writing; non-native English speakers may use LLMs differently for translation/proofreading; Database variation: OpenAlex uses first publication date (often preprints) rather than formal publication date, appearing about 1 year ahead of other databases; Selection bias: Analysis limited to six databases; may overlook other academic databases or non-indexed publications; Language bias: Analysis restricted to English-language terms; non-English papers with translated English abstracts may introduce artifacts; Temporal bias: Normalization by publication volume helps mitigate, but absolute growth in publication output may confound results; Disciplinary bias: Different fields have varying rates of LLM adoption; STEM fields show much higher uptake than social sciences/humanities; Open-access bias: PMC analysis covers only open-access biomedical/life science publications, excluding paywalled journals like The Lancet, JAMA; Term selection bias: 12 terms chosen based on prior literature and exploratory analysis; other LLM-associated terms may exist but were not included; Causality assumption: Study cannot definitively prove LLMs cause increased term usage; could reflect editorial changes, author preferences, or other factors; Selection bias: Only open-access PMC papers analyzed, excluding many leading medical journals (The New England Journal of Medicine, The Lancet, JAMA); Language bias: Analysis limited to English-language terms and abstracts; non-English articles with translated English abstracts could artificially inflate LLM-term counts; Discipline-specific bias: Some terms naturally occur in certain disciplines (e.g., 'foster' in education, 'heighten' in psychology) independent of LLM use; Temporal confounding: Study cannot distinguish whether increases reflect direct LLM text generation, editorial/publishing practice changes linked to LLM use, or longer-term trends unrelated to LLMs; Database heterogeneity: Different databases use different indexing practices; OpenAlex uses first publication date (often preprint) rather than formal publication date
Limitations
- "The results are limited by the set of 12 LLM-associated terms used and the six databases analysed, and we may have overlooked other terms commonly suggested by LLMs, including those used by models other than ChatGPT." Additionally, "some of the studied terms may naturally be used more frequently in specific disciplines, independent of LLM usage
- For example, the term foster or fostering may commonly appear in fields like education, social work, child development or animal care, making these counts less reliable as direct indicators of LLM influence." Furthermore, "The PMC full-text dataset is available only for open-access biomedical and life-science research, and not all articles are included in the analysis
- As a result, the findings mainly reflect LLM-term usage in open-access publications and may not fully represent non-open-access journals." The authors also note that "this study analysed changes in the frequency and co-occurrence of selected LLM-associated words in academic publications but did not assess potential semantic or stylistic changes in sentence structure which may also be influenced by LLMs."
Open questions raised
- How different LLMs (ChatGPT, Gemini, DeepSeek) differ in their influence on academic writing
- Whether patterns reflect direct LLM text generation or indirect changes in editorial/publishing practices
- Language patterns in non-English academic publications
- Semantic and stylistic changes in sentence structure beyond word frequency
- Analysis of a broader range of alternative terms beyond those tested
- Investigation of whether LLM influence differs across specific subject disciplines
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations