Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical Knowledge
Annette Flanagin, Kirsten Bibbins‐Domingo, Michael Berkwits, Stacy Christiansen · JAMA · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1001/jama.2023.1344
Methodology & findings
Study design
Editorial commentary and policy analysis.
Main result
The editorial documents that "In January 2023, Nature reported on 2 preprints and 2 articles published in the science and health fields that included ChatGPT as a bylined author" and notes that "these articles and their nonhuman "authors" have already been indexed in PubMed and Google Scholar." The authors found that experiments with ChatGPT showed "text responses to questions, while mostly well written, are formulaic (which was not easily discernible), not up to date, false or fabricated, without accurate or complete references, and worse, with concocted nonexistent evidence for claims or statements it makes."
Reports effect sizes.
Research paradigm
Critical interpretivism; normative/regulatory analysis
Author conclusions
The authors conclude that "In this era of pervasive misinformation and mistrust, responsible use of AI language models and transparent reporting of how these tools are used in the creation of information and publication are vital to promote and protect the credibility and integrity of medical research and trust in medical knowledge." They also state that "AI technologies have existed for some time, will be further and faster developed, and will continue to be used in all stages of research and the dissemination of information, hopefully with innovative advances that offset any perils."
Open questions raised
- The authors identify that the publishing community must address evolving risks and opportunities as "Transformative, disruptive technologies, like AI language models, create promise and opportunities as well as risks and threats for all involved in the scientific enterprise." They note that "Calls for journals to implement screening for AI-generated content will likely escalate," and that future EQUATOR Network reporting guidelines are under development for prognostic and diagnostic studies using AI and machine learning (STARD-AI and TRIPOD-AI).
- The authors identify that the EQUATOR Network has "several other reporting guidelines in development for prognostic and diagnostic studies that use AI and machine learning, such as STARD-AI and TRIPOD-AI," suggesting ongoing work to address gaps in guidance for AI use in research.
- The authors identify the need for screening tools to detect AI-generated content and note that "Calls for journals to implement screening for AI-generated content will likely escalate," though they acknowledge that "with large investments in further development, AI tools may be capable of evading any such screens."
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations