Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers
Catherine A. Gao, Frederick M. Howard, Nikolay S. Markov, Emma Dyer, Siddhi Ramesh, Yuan Luo et al. · npj Digital Medicine · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41746-023-00819-6
Methodology & findings
Study design
Experimental comparison study with human evaluation and detector tool assessment.
Sample
N = 500, 3 groups
Primary method
Descriptive statistics including medians and interquartile ranges; AUROC calculation for detector performance; plagiarism detection via automated tools (iThenticate and plagiarism detector website); blinded human review with accuracy assessment.
Main result
The study found that most generated abstracts were detected using an AI output detector with "median [interquartile range] of 99.98% 'fake' [12.73%, 99.98%]" compared to "median 0.02% [IQR 0.02%, 0.09%] for the original abstracts." Additionally, "when given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/positivist
Author conclusions
The authors conclude that "ChatGPT writes believable scientific abstracts, though with completely generated data." They further state that "Depending on publisher-specific guidelines, AI output detectors may serve as an editorial tool to help maintain scientific standards."
Risk of bias
Selection bias: Only abstracts from five high-impact factor medical journals were included, which may not represent all scientific domains or journals with different impact factors; Potential detector bias: Reliance on GPT-2 Output Detector, which may have inherent biases in classification; Human reviewer bias: Blinded reviewers' judgments may be influenced by writing style patterns specific to ChatGPT that differ from journal to journal; Limited sample size for human review: Not explicitly stated in abstract how many reviewers participated or how many abstracts were evaluated per reviewer; Selection bias: Abstracts gathered from only five high-impact medical journals may not represent broader journal landscape; Blinding: While human reviewers were blinded, the selection of journals and abstracts could introduce bias; Potential confounding: Journal prestige and field-specific writing conventions not controlled; Selection bias: abstracts gathered from five high-impact journals may not represent full spectrum of scientific writing; Potential human reviewer bias in blinded assessment despite blinding procedures; Limited specification of journal selection criteria; Detector performance may vary across different types of scientific abstracts
Open questions raised
- The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," indicating need for further research on policy development and standardization across publishing venues.
- The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," suggesting need for clearer guidelines and continued investigation.
- The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," indicating need for further work on policy development and guidelines.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations