12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Delving into the Utilisation of ChatGPT in Scientific Publications in Astronomy

Simone Astarita, Sandor Kruk, Jan Reerink, Pablo Gómez · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
2
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2406.17324

Methodology & findings

Study design

Quantitative linguistic analysis using word frequency analysis and statistical hypothesis testing.

Sample

N = 1061637, 6 groups

Primary method

Kolmogorov–Smirnov test (two-sample, non-parametric test for comparing distributions); Tokenization using spaCy English model; Frequency ratio calculation: |w ∈ Cai|/|Cai| : |w ∈ Chum|/|Chum|; Percentage change calculation year-over-year; Averaging of percentages across word subsets (top 5, 10, 20, 50, and all 100); Normalization by annual publication count

Main result

The study found that "the frequency of occurrence of words in W increased significantly in 2024, especially for the words near the top, and does so to a higher extent when compared to other control words (C) or years." The analysis identified ChatGPT-favoured words such as "delve" and "underscore" appearing in a much larger portion of astronomy papers in 2024 compared to previous years, with all p-values under 0.05 in the Kolmogorov-Smirnov test comparing 2024 to other years.

Reports effect sizes.

Research paradigm

Positivist/Empiricist

Author conclusions

"We conclude that the frequency distribution of words in W is significantly different in 2024 compared to all years before 2023, and this variation is unlikely due to random language fluctuations." The authors state: "our study provides evidence that: ChatGPT-generated text overuses the words in W, such as delve and underscore, when compared to humans" and "In astronomical publications, the frequency of occurrence of words in W increased significantly in 2024, especially for the words near the top." They encourage that "organisations, publishers, and researchers work together to establish ethical and pragmatic guidelines to maximise the benefits of these systems while maintaining scientific rigour."

Risk of bias

Confounding by author native language distribution over time; Lack of control for paper length changes across years; Binary word presence detection (not accounting for word frequency within papers); Potential seasonal publication bias in timing of dataset extraction (24 May 2024); Selection bias toward English-language astronomy publications indexed by NASA ADS; Inability to distinguish ChatGPT use from other LLMs; Confounding from changes in author demographic composition (native language distribution); Lack of control for average paper length changes over time; Binary treatment of word presence (presence/absence only, not frequency within documents); Scope limited to ChatGPT; other LLM impacts not assessed; Potential temporal bias from growing publication rates in astronomy; Confounding variables: unable to control for changes in distribution of authors' native languages; Not accounting for average paper length changes over time; Selection bias in word identification: relies on external corpus (Liang et al.) created from specific ChatGPT version (3.5); Scope limitation: only ChatGPT detection, ignoring other LLMs; Binary presence/absence metric: does not quantify word frequency within papers, only presence across papers

Limitations

  • "The presented data and analysis have some limitations: a study of the temporal evolution of scientific writing is faced with many confounders
  • We were unable to control for changes due to differences in the distribution of authors' native languages, which are known to impact language
  • Our analysis does not directly consider the average paper length changes over time, which can impact the frequency of occurrence
  • however, we hope to have mitigated this effect by comparing our results with a control group of words
  • We also considered only whether a word is present or not in the text, not the number of times it appears." Additionally, "our scope is limited to identifying ChatGPT-related characteristics given its high adoption rate
  • the impact of other LLMs is hence neglected."

Open questions raised

  • Need for direct estimation of proportion of AI-generated text in published works
  • Assessment of impact of other LLMs beyond ChatGPT
  • Investigation of whether ChatGPT-favoured words indicate actual ChatGPT usage vs. coincidental language patterns
  • Study of mechanisms by which LLMs could deteriorate model quality through training on AI-generated data
  • Evaluation of reliability of AI-generated text detection tools
  • Research on comparative quality of AI-generated vs human-written scientific content
Data: NASA Astrophysics Data System (ADS) astronomy collection articles (1 January 2000 - 24 May 2024); Liang et al. corpus of ChatGPT-generated vs human-written academic text (https://arxiv.org/abs/2404.01268v1); NASA Astrophysics Data System (ADS) - 1,061,637 astronomy articles from 2000-2024; Liang et al. corpus of ChatGPT 3.5 vs. human-authored academic text paragraphs; NASA Astrophysics Data System (ADS) - 1,061,637 astronomy articles from 2000-2024 (accessed via API on 24 May 2024); Liang et al. corpus of human and AI-generated academic text (https://arxiv.org/abs/2404.01268v1)Code: https://github.com/ESA-Datalabs/llm-usage-astronomy/Extracted from: pdfAgreement 54%

Explore related topics

Related papers