What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience
Filip Miletić, Neele Falk · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.63317/3ai7wig4fhd8
Methodology & findings
Study design
Mixed-methods study combining: (1) diachronic corpus analysis of 37,760 NLP papers from ACL Anthology comparing two time periods (2020-2022 vs.
Sample
N = 37760, 15 groups
Primary method
Log-likelihood score (Rayson and Garside, 2000) for identifying word frequency changes; Spearman correlation for analyzing relationship between log-likelihood and neighborhood density; word2vec (skip-gram with negative sampling, window size 5, 100 dimensions) trained separately for t1 and t2 with three runs for robustness; Mann-Whitney-U test (0.05 level) for significance testing on neighborhood density changes; k-means clustering (k=8, scikit-learn) on ModernBERT contextualized embeddings; logistic regression with elastic net regularization (stability selection, 5-fold inner and 10-fold outer cross-validation); generalized mixed-effects model with paper ID as random effect for LLM dataset analysis; BERTopic for topic analysis; feature extraction via elfen and LFTK; standardization and scaling of ~1,000 linguistic features; correlation analysis (retaining features with r<0.7) for feature selection; 10-fold cross-validation for unbiased estimates
Main result
The study found that "post-ChatGPT papers are distinguished by more complex lexical choices (e.g., enhance rather than improve) and further stylistic properties (e.g., lower lexical diversity)." Additionally, "LLM-improved texts are perceived as clearer and more exciting," with "the strongest preference for LLM-improved texts can be observed in clarity and excitement."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed methods (quantitative computational analysis + qualitative human annotation)
Author conclusions
"In this work, we investigated the influence of LLMs on writing style in scientific communication, specifically in the NLP domain. With regard to word usage, we found that both word frequency and the contexts in which these are used have changed significantly, indicating semantic specialization in some cases and generalization in others." The authors further state: "Crucially, trends in word usage as well as stylistic features are broadly consistent across the naturalistic and synthetic data scenarios, indicating that many observed shifts in writing practices can be attributed to growing LLM use. Finally, in a pilot study, we measured the impact of these changes in writing style on the reading experience. The results indicate that LLM-improved texts are perceived as more understandable and exciting."
Risk of bias
Selection bias: Binary time period comparison may oversimplify temporal trends; gap between t1 and t2 may miss intermediate changes; Confounding variables in naturalistic corpus: Other AI-assisted writing tools (Grammarly) may have influenced t1 period; topical shifts in NLP research could confound stylistic changes; Annotation bias: Small sample (20 domain experts) with only 2 judgments per instance; participants reported making assumptions about which text was LLM-based; acknowledged mismatch between negative attitudes toward LLM writing and positive ratings; Measurement bias: LLM dataset uses only GPT-3.5-turbo; later papers may have used different LLMs (Claude) not represented in synthetic data; Coverage issues in original ACL-OCL corpus affecting representativeness; Selection bias: Study limited to NLP papers from ACL Anthology, may not generalize to other scientific domains; Temporal confounding: Changes between t1 and t2 could reflect topic shifts beyond LLM use (though topic analysis was performed); Annotator bias: Only 20 domain experts; qualitative interviews revealed subjective variation in authenticity and trustworthiness ratings; Model-specific bias: Used only GPT-3.5-turbo for synthetic dataset generation, may not represent other LLMs (Claude, GPT-4, etc.); Reference removal bias: Authors note LLMs remove references when improving text, confounding text quality measures; Publication bias: ACL papers may represent high-quality writing already, not representative of broader scientific writing; Selection bias in corpus: papers may have differential coverage across time periods due to crawling issues; temporal confounding: topical shifts within NLP research between periods could explain some linguistic changes independent of LLM use; annotation bias: 20 domain expert raters may have biases regarding LLM-generated text; detection bias: participants attempted to identify which text was LLM-generated, potentially introducing motivated reasoning; generalization bias: findings limited to NLP domain in English only.
Limitations
- "First, our analysis is framed as a binary comparison of language use across two time periods
- This approach has important practical benefits and is underpinned by clearly stated assumptions (e.g., uncertainty regarding the precise degree of LLM use in the second time period)
- However, a finer-grained comparison of smaller time slices could provide a clearer picture regarding the development patterns of LLM-supported stylistic choices." Additionally, "our study is limited to a single scientific domain (Natural Language Processing) and one European language (English)." The authors also note: "Finally, we ran a pilot annotation study with 20 participants, collecting two judgments per instance
- A larger-scale setup with more annotations per instance would provide more robust results."
Open questions raised
- Finer-grained comparison of smaller time slices for clearer picture of stylistic development patterns
- Larger empirical study comparing stylistic differences between traditional writing tools (Grammaly) and modern LLMs
- Replication across multiple scientific domains to assess generalizability
- Analysis in languages other than English
- Larger-scale annotation study with more participants and multiple judgments per instance to verify generalizability of reading experience results
- Exploration of prompt-specific and model-specific characteristics of writing style and word usage
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations