Delving into the Utilisation of ChatGPT in Scientific Publications in Astronomy
Simone Astarita, Sandor Kruk, Jan Reerink, Pablo Gómez · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2406.17324
Methodology & findings
Study design
Quantitative linguistic analysis using word frequency analysis and statistical hypothesis testing.
Sample
N = 1061637, 6 groups
Primary method
Kolmogorov–Smirnov test (two-sample, non-parametric test for comparing distributions); Tokenization using spaCy English model; Frequency ratio calculation: |w ∈ Cai|/|Cai| : |w ∈ Chum|/|Chum|; Percentage change calculation year-over-year; Averaging of percentages across word subsets (top 5, 10, 20, 50, and all 100); Normalization by annual publication count
Main result
The study found that "the frequency of occurrence of words in W increased significantly in 2024, especially for the words near the top, and does so to a higher extent when compared to other control words (C) or years." The analysis identified ChatGPT-favoured words such as "delve" and "underscore" appearing in a much larger portion of astronomy papers in 2024 compared to previous years, with all p-values under 0.05 in the Kolmogorov-Smirnov test comparing 2024 to other years.
Reports effect sizes.
Research paradigm
Positivist/Empiricist
Author conclusions
"We conclude that the frequency distribution of words in W is significantly different in 2024 compared to all years before 2023, and this variation is unlikely due to random language fluctuations." The authors state: "our study provides evidence that: ChatGPT-generated text overuses the words in W, such as delve and underscore, when compared to humans" and "In astronomical publications, the frequency of occurrence of words in W increased significantly in 2024, especially for the words near the top." They encourage that "organisations, publishers, and researchers work together to establish ethical and pragmatic guidelines to maximise the benefits of these systems while maintaining scientific rigour."
Risk of bias
Confounding by author native language distribution over time; Lack of control for paper length changes across years; Binary word presence detection (not accounting for word frequency within papers); Potential seasonal publication bias in timing of dataset extraction (24 May 2024); Selection bias toward English-language astronomy publications indexed by NASA ADS; Inability to distinguish ChatGPT use from other LLMs; Confounding from changes in author demographic composition (native language distribution); Lack of control for average paper length changes over time; Binary treatment of word presence (presence/absence only, not frequency within documents); Scope limited to ChatGPT; other LLM impacts not assessed; Potential temporal bias from growing publication rates in astronomy; Confounding variables: unable to control for changes in distribution of authors' native languages; Not accounting for average paper length changes over time; Selection bias in word identification: relies on external corpus (Liang et al.) created from specific ChatGPT version (3.5); Scope limitation: only ChatGPT detection, ignoring other LLMs; Binary presence/absence metric: does not quantify word frequency within papers, only presence across papers
Limitations
- "The presented data and analysis have some limitations: a study of the temporal evolution of scientific writing is faced with many confounders
- We were unable to control for changes due to differences in the distribution of authors' native languages, which are known to impact language
- Our analysis does not directly consider the average paper length changes over time, which can impact the frequency of occurrence
- however, we hope to have mitigated this effect by comparing our results with a control group of words
- We also considered only whether a word is present or not in the text, not the number of times it appears." Additionally, "our scope is limited to identifying ChatGPT-related characteristics given its high adoption rate
- the impact of other LLMs is hence neglected."
Open questions raised
- Need for direct estimation of proportion of AI-generated text in published works
- Assessment of impact of other LLMs beyond ChatGPT
- Investigation of whether ChatGPT-favoured words indicate actual ChatGPT usage vs. coincidental language patterns
- Study of mechanisms by which LLMs could deteriorate model quality through training on AI-generated data
- Evaluation of reliability of AI-generated text detection tools
- Research on comparative quality of AI-generated vs human-written scientific content
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations