12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Have Large Language Models Enhanced the Way Civil & Environmental Engineers Write? A Quantitative Analysis of Scholarly Communication over 25 Years

Morgan D. Sanger, Brett W. Maurer · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Corpus linguistic analysis using frequency-shift methodology on 149,452 abstracts from American Society of Civil Engineers publications (2000-2025).

Sample

N = 149452, 10 groups

Primary method

Least-squares linear regression for counterfactual trendline fitting to pre-2022 vocabulary frequencies. Frequency gap (δ) and frequency ratio (r) calculations for deviation detection. Threshold tuning via iterative optimization to maximize group-prevalence differentials (ΔGrare and ΔGcommon). Sensitivity analysis to evaluate false-positive rates and positive prediction rates. Document-level frequency calculations normalized by publication year. All implemented in Python with Natural Language Toolkit (NLTK).

Main result

Beginning around 2023, the frequencies of many stylistic marker words (e.g., enhance) sharply depart from historical trajectories. "Abstracts classified as likely LLM-assisted exhibit increased lexical diversity, comma use, and complexity, with reduced passive voice and hedging language, producing prose that is more segmented, complex, and confident." Using the frequency-shift methodology, estimates of LLM-associated linguistic patterns are 15.3% in 2024 and 26.2% in 2025, with domain-specific rates as high as 38.4% in construction engineering.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empirical quantitative

Author conclusions

"Several clear conclusions emerge regarding how LLM use is reshaping research prose. Prior to the widespread adoption of LLMs, CEE scholarship exhibits long-term trends toward increasing numbers of authors, longer abstracts and sentences, greater use of segmenting punctuation, higher required reading levels, and a shift toward active, first-person verb constructions." "Abstracts classified as likely LLM-written exhibit systematic shifts, including increased word choice diversity, heavier use of commas, reduced readability, decreased reliance on passive constructions, and diminished use of qualifying language used to convey uncertainty. These features collectively suggest prose that is more segmented, more syntactically complex, and more assertive, such that LLM adoption is measurably impacting the stylistic evolution of scholarly writing in civil and environmental engineering."

Risk of bias

Selection bias: Study limited to ASCE publications only; may not generalize to other publishers or disciplines; Classification bias: Binary classification of abstracts as LLM-written based on marker word thresholds lacks ground truth validation; ~1% false-positive rate in pre-2023 period; Measurement bias: Proxy metrics for voice (first-person pronouns for active voice, past participles for passive voice) are incomplete assessments; Temporal bias: Publication lag between writing and publication may obscure when LLM adoption actually began; Confounding: Editorial practices, evolving disciplinary rhetoric, and emergent research topics may contribute to vocabulary shifts independent of LLM use; No ground-truth validation dataset for LLM-written abstracts; Reliance on proxy metrics for voice and confidence (first-person pronouns, passive voice, hedging words); Potential temporal bias in editor practices and publication standards; Limited generalizability beyond ASCE journals; Potential confounding from other contemporaneous linguistic trends; No ground-truth validation dataset for LLM classification; Potential publisher text encoding artifacts affecting punctuation detection; Classification thresholds based on sensitivity analysis may not generalize across different LLM versions or prompts; Potential confounding by editorial changes, evolving disciplinary rhetoric, or shifts in publication practices; Pre-2023 transition year shows low LLM prevalence, potentially confounding trend detection; Style/content word classification involves subjective judgment

Limitations

  • "Most notably, no ground-truth dataset exists identifying which abstracts were written with or without LLM assistance
  • Our methodology thus focuses on detecting population-level linguistic signals
  • these signals are used to make inferences about LLM use that are necessarily uncertain." Additionally, "Abstracts are also a particularly efficient entry point for LLM assistance because of their summarization and synthesis functions
  • Thus, the classification of an abstract as likely LLM-written does not imply LLM authorship of the manuscript." The authors note limitations in readability metrics: "their oversimplification of complexity in terms of syllable quantity and sentence length
  • and (iii) their lack of consideration for context and reader familiarity with the subject matter." Mean values are used rather than full distributional analysis, and linear counterfactual trendlines may not capture non-linear linguistic evolution.

Open questions raised

  • Lack of ground-truth datasets identifying which abstracts were definitively written with or without LLM assistance
  • Unknown whether LLM classification findings generalize beyond ASCE publications to other publishers and disciplines
  • No examination of full manuscript texts; only abstracts analyzed due to public availability constraints
  • Alternative nonlinear models for counterfactual trendlines could be developed in future work
  • Refinement of style versus content word classifications needed
  • Need for more rigorous statistical analysis of full distributions rather than mean values
Data: Publicly available abstracts from American Society of Civil Engineers (ASCE) journals and conference proceedings (2000-2025). Data source: ASCE publications across 30 journals and 935 conference proceedings. Total corpus: 149,452 abstracts. Specific dataset availability not explicitly stated as open-access beyond public availability of original abstracts.; 149,452 abstracts from ASCE journals and conference proceedings (2000-2025) - publicly available through ASCE; 149,452 ASCE abstracts from 2000-2025 (publicly available through ASCE journals and conference proceedings)Extracted from: pdfAgreement 54%

Explore related topics

Related papers