12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Delving into LLM-assisted writing in biomedical publications through excess vocabulary

Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, Jan Lause · Science Advances · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
90
Citations
37.45
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1126/sciadv.adt3813

Methodology & findings

Study design

Large-scale observational analysis of vocabulary changes in biomedical abstracts indexed by PubMed (2010-2024).

Sample

N = 15000000, 3 groups

Primary method

Excess vocabulary analysis (statistical methodology for detecting anomalous word frequency patterns); temporal trend analysis of word frequency changes from 2010 to 2024 across PubMed abstracts.

Main result

The study found that "at least 13.5% of 2024 abstracts were processed with LLMs" based on excess vocabulary analysis of more than 15 million biomedical abstracts from 2010 to 2024, with this lower bound "differing across disciplines, countries, and journals, reaching 40% for some subcorpora." The authors demonstrate that "LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the COVID pandemic."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative analysis of linguistic patterns in large-scale dataset

Author conclusions

The authors conclude that "LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the COVID pandemic," and that the "appearance of LLMs led to an abrupt increase in the frequency of certain style words," with the excess word analysis providing evidence that "at least 13.5% of 2024 abstracts were processed with LLMs."

Risk of bias

The excess vocabulary approach may not capture all LLM usage patterns, potentially underestimating prevalence; The method relies on detecting certain 'style words' as markers of LLM use, which could be subject to false positives/negatives; Cross-disciplinary, cross-country, and cross-journal variations in LLM adoption patterns may confound results; The excess word analysis may not capture all forms of LLM usage or may misattribute natural linguistic evolution to LLM influence; Variation across disciplines, countries, and journals suggests potential confounding by publication practices and regional differences in LLM adoption; Selection bias: Analysis limited to PubMed-indexed abstracts, which may not represent all biomedical research; Confounding: Other factors beyond LLM adoption could drive vocabulary changes; Detection bias: Excess vocabulary markers may correlate with but not definitively prove LLM usage; Geographic/disciplinary bias: Unequal distribution of LLM usage across subcorpora suggests differential adoption patterns

Data: not_statedCode: not_statedExtracted from: pdfAgreement 55%

Explore related topics

Related papers