‘As of my last knowledge update’: How is content generated by ChatGPT infiltrating scientific papers published in premier journals?
Artur Strzelecki · Learned Publishing · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/leap.1650
Methodology & findings
Study design
Systematic scoping review using SPAR4SLR framework adapted for identifying AI-generated content.
Sample
N = 1362, 9 groups
Primary method
Descriptive bibliometric analysis. Manual content analysis of full texts. Keyword-based discipline classification using lexicon of discipline-specific keywords. Frequency counts and proportional analysis. Citation tracking via Google Scholar. No inferential statistical testing conducted.
Main result
The study identified 1,362 scientific articles containing unequivocal ChatGPT-generated content through systematic Google Scholar searches. Key findings revealed: "in highly regarded premier journals, articles appear that bear the hallmarks of the content generated by AI large language models, whose use was not declared by the authors (1); many of these identified papers are already receiving citations from other scientific works, also placed in journals found in scientific databases (2); and, most of the identified papers belong to the disciplines of medicine and computer science, but there are also articles that belong to disciplines such as environmental science, engineering, sociology, education, economics and management (3)." Of the 89 articles identified in indexed journals, 64 were published in Q1 journals, and 60 articles had already been cited with a total of 528 citations.
Reports effect sizes.
Research paradigm
Empirical-positivist (detection and classification of AI-generated text artifacts)
Author conclusions
"AI will inevitably have an impact on scientific research in the future, from copyediting to writing literature reviews. It is important to be open and transparent about the extent to which AI has been and will continue to be used, particularly in publications and scientific research." The author also concludes that "the presence of these papers underscores a fundamental systemic concern related to the publish-or-perish pressure experienced by academics, frequently compounded by time constraints and resource limitations" and calls for "ongoing monitoring, prompting institutions, funding bodies, peer reviewers and researchers to engage in critical self-reflection."
Risk of bias
Selection bias: Study relies on identification of specific ChatGPT phrases; AI-generated content without distinctive phrases would be missed; Detection bias: Perfect string matching may not capture paraphrased or edited AI-generated content; Indexing bias: Focus on Google Scholar and Web of Science/Scopus may exclude papers in non-indexed journals; Temporal bias: Journal indexing status changes over time; some papers may have been later retracted or removed; Selection bias: Only articles containing specific ChatGPT phrases were identified; papers with heavily edited or modified AI-generated content may be missed; String-matching bias: Reliance on perfect phrase matching limits detection capability; Database bias: Google Scholar indexing speed may introduce temporal bias compared to Web of Science/Scopus; Confirmation bias: Manual analysis introduces subjective interpretation of whether content is AI-generated; Selection bias: Restriction to papers with identifiable ChatGPT 'signature' phrases may exclude modified AI content without telltale markers; Detection bias: Perfect string matching method may miss AI-generated content that authors edited to remove characteristic phrases; Database indexing bias: Google Scholar faster than Web of Science/Scopus, creating temporal lag in indexing; Discipline bias: Over-representation in medicine and computer science may reflect differential adoption of AI tools or differential detection capability; Publication bias: Focus on indexed journals excludes potentially higher prevalence in lower-tier journals
Limitations
- The author stated: "This study is not without limitations
- First, some papers containing AI-generated text are not yet indexed in scientific databases, because Google Scholar is faster and content was primarily identified in Google Scholar
- Second, the identification of research papers featuring evident AI-generated text relies solely on perfect string matching, which has its limitations
- More advanced methodologies for detecting text likely to be significantly influenced or generated by AI models are discussed by Liang et al
- (2024) and Yeadon et al
- Third, there may be many cases of AI content being used, but without the errors and without the typical phrases, in cases where the author modified it a bit
Open questions raised
- The paper identifies that existing research has not adequately addressed how to identify papers partially generated by ChatGPT or other LLMs beyond linguistic corrections. The gap focuses on detecting content creation by AI tools rather than language editing. Future research directions implied include: developing more advanced detection methodologies for AI-generated content, evaluating publisher response mechanisms, reassessing publication frameworks to prevent erroneous content, and exploring strategies to cultivate integrity in scientific publishing.
- The paper identifies a research gap: "Given what has been discussed so far, a newly identified research gap is how to find and recognize scientific articles that have been partially generated by ChatGPT or any other LLM." Future research directions include: developing more advanced detection methodologies beyond string matching, monitoring evolution of LLM capabilities and detection evasion techniques, and reassessing publication frameworks to address systemic flaws in peer review and editorial processes.
- The paper identifies the following research gaps and future directions: (1) Need for more advanced methodologies beyond perfect string matching for detecting AI-influenced or generated text; (2) Necessity for systematic reassessment and restructuring of publication frameworks to prevent erroneous content promotion; (3) Requirement for ongoing monitoring and critical self-reflection by institutions, funding bodies, peer reviewers, and researchers; (4) Exploration of strategies to cultivate integrity culture in scientific publishing and implementation of repercussions for breaches; (5) Need to establish clearer, more consistent publisher policies distinguishing between different types and levels of AI tool use across all major publishers.
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations