12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience

Filip Miletić, Neele Falk · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.63317/3ai7wig4fhd8

Methodology & findings

Study design

Mixed-methods study combining: (1) diachronic corpus analysis of 37,760 NLP papers from ACL Anthology comparing two time periods (2020-2022 vs.

Sample

N = 37760, 15 groups

Primary method

Log-likelihood score (Rayson and Garside, 2000) for identifying word frequency changes; Spearman correlation for analyzing relationship between log-likelihood and neighborhood density; word2vec (skip-gram with negative sampling, window size 5, 100 dimensions) trained separately for t1 and t2 with three runs for robustness; Mann-Whitney-U test (0.05 level) for significance testing on neighborhood density changes; k-means clustering (k=8, scikit-learn) on ModernBERT contextualized embeddings; logistic regression with elastic net regularization (stability selection, 5-fold inner and 10-fold outer cross-validation); generalized mixed-effects model with paper ID as random effect for LLM dataset analysis; BERTopic for topic analysis; feature extraction via elfen and LFTK; standardization and scaling of ~1,000 linguistic features; correlation analysis (retaining features with r<0.7) for feature selection; 10-fold cross-validation for unbiased estimates

Main result

The study found that "post-ChatGPT papers are distinguished by more complex lexical choices (e.g., enhance rather than improve) and further stylistic properties (e.g., lower lexical diversity)." Additionally, "LLM-improved texts are perceived as clearer and more exciting," with "the strongest preference for LLM-improved texts can be observed in clarity and excitement."

Reports effect sizes and confidence intervals.

Research paradigm

Mixed methods (quantitative computational analysis + qualitative human annotation)

Author conclusions

"In this work, we investigated the influence of LLMs on writing style in scientific communication, specifically in the NLP domain. With regard to word usage, we found that both word frequency and the contexts in which these are used have changed significantly, indicating semantic specialization in some cases and generalization in others." The authors further state: "Crucially, trends in word usage as well as stylistic features are broadly consistent across the naturalistic and synthetic data scenarios, indicating that many observed shifts in writing practices can be attributed to growing LLM use. Finally, in a pilot study, we measured the impact of these changes in writing style on the reading experience. The results indicate that LLM-improved texts are perceived as more understandable and exciting."

Risk of bias

Selection bias: Binary time period comparison may oversimplify temporal trends; gap between t1 and t2 may miss intermediate changes; Confounding variables in naturalistic corpus: Other AI-assisted writing tools (Grammarly) may have influenced t1 period; topical shifts in NLP research could confound stylistic changes; Annotation bias: Small sample (20 domain experts) with only 2 judgments per instance; participants reported making assumptions about which text was LLM-based; acknowledged mismatch between negative attitudes toward LLM writing and positive ratings; Measurement bias: LLM dataset uses only GPT-3.5-turbo; later papers may have used different LLMs (Claude) not represented in synthetic data; Coverage issues in original ACL-OCL corpus affecting representativeness; Selection bias: Study limited to NLP papers from ACL Anthology, may not generalize to other scientific domains; Temporal confounding: Changes between t1 and t2 could reflect topic shifts beyond LLM use (though topic analysis was performed); Annotator bias: Only 20 domain experts; qualitative interviews revealed subjective variation in authenticity and trustworthiness ratings; Model-specific bias: Used only GPT-3.5-turbo for synthetic dataset generation, may not represent other LLMs (Claude, GPT-4, etc.); Reference removal bias: Authors note LLMs remove references when improving text, confounding text quality measures; Publication bias: ACL papers may represent high-quality writing already, not representative of broader scientific writing; Selection bias in corpus: papers may have differential coverage across time periods due to crawling issues; temporal confounding: topical shifts within NLP research between periods could explain some linguistic changes independent of LLM use; annotation bias: 20 domain expert raters may have biases regarding LLM-generated text; detection bias: participants attempted to identify which text was LLM-generated, potentially introducing motivated reasoning; generalization bias: findings limited to NLP domain in English only.

Limitations

  • "First, our analysis is framed as a binary comparison of language use across two time periods
  • This approach has important practical benefits and is underpinned by clearly stated assumptions (e.g., uncertainty regarding the precise degree of LLM use in the second time period)
  • However, a finer-grained comparison of smaller time slices could provide a clearer picture regarding the development patterns of LLM-supported stylistic choices." Additionally, "our study is limited to a single scientific domain (Natural Language Processing) and one European language (English)." The authors also note: "Finally, we ran a pilot annotation study with 20 participants, collecting two judgments per instance
  • A larger-scale setup with more annotations per instance would provide more robust results."

Open questions raised

  • Finer-grained comparison of smaller time slices for clearer picture of stylistic development patterns
  • Larger empirical study comparing stylistic differences between traditional writing tools (Grammaly) and modern LLMs
  • Replication across multiple scientific domains to assess generalizability
  • Analysis in languages other than English
  • Larger-scale annotation study with more participants and multiple judgments per instance to verify generalizability of reading experience results
  • Exploration of prompt-specific and model-specific characteristics of writing style and word usage
Data: Updated ACL-OCL corpus (99.2k papers): https://github.com/FilipMiletic/ScientificCommunication; 3,000 pairs of human-written texts and LLM-produced improvements with annotations of human reading experience for 200 pairs: https://github.com/FilipMiletic/ScientificCommunication; Updated ACL-OCL corpus: 99.2k papers from ACL Anthology (published until end of 2024); Synthetic LLM dataset: 3,000 pairs of human-written texts and their LLM-produced improvements; Annotation data: 200 annotated text pairs with human reading experience ratings; Repository: https://github.com/FilipMiletic/ScientificCommunication; ACL-OCL corpus (updated): 99.2k papers from ACL Anthology published until end of 2024, available with code for future updates; LLM dataset: 3,000 pairs of human-written texts and their LLM-produced improvementsCode: GitHub: https://github.com/FilipMiletic/ScientificCommunication (contains corpus update pipeline, one-line script for ingesting future papers, annotation data); https://github.com/FilipMiletic/ScientificCommunication - contains full configuration, annotation dataset, and annotation results; One-line command script provided for ingesting future papers into updated corpus pipeline; https://github.com/FilipMiletic/ScientificCommunication (annotation code, dataset, and results)Extracted from: pdfAgreement 45%

Explore related topics

Related papers