AI vs academia: Experimental study on AI text detectors’ accuracy in behavioral health academic writing
Andrey A Popkov, Tyson S. Barrett · Accountability in Research · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1080/08989621.2024.2331757
Methodology & findings
Study design
Experimental study using content analysis of academic texts and AI-generated texts.
Sample
N = 300, 9 groups
Main result
The study found that "The free AI detector showed a median of 27.2% for the proportion of academic text identified as AI-generated, while commercial software Originality.AI demonstrated better performance but still had limitations, especially in detecting texts generated by Claude." These results indicate problematic false positive and false negative rates in both free and paid AI detection tools when applied to behavioral health academic writing.
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that "These error rates raise doubts about relying on AI detectors to enforce strict policies around AI text generation in behavioral health publications." They also state that "the implementation of such policies requires accurate AI detection tools" and that "Inaccurate detectors risk unnecessary penalties for human authors and/or may compromise the effective enforcement of guidelines against AI-generated content."
Risk of bias
Selection bias: Sample limited to behavioral health publications from 2016-2018; Potential detector-specific bias: Only two detectors tested (one free, Originality.AI); Model-specific variation: Different AI models (ChatGPT vs Claude) showed differential detection rates; Selection bias in article selection from behavioral health and psychiatry journals; Temporal limitation: articles from 2016-2018 may not represent current publication styles; Potential operator bias in how texts were selected or processed; Limited geographic/journal diversity information not provided in abstract; Selection bias: only behavioral health publications sampled (2016-2018); Potential detector bias: only two AI text generators tested (ChatGPT and Claude); Temporal bias: older publications (2016-2018) may not reflect current academic writing styles; Limited generalizability: results may not apply to other academic disciplines
Open questions raised
- The authors identify that the accuracy of AI text detection tools in identifying human-written versus AI-generated content has been found to vary across published studies, suggesting a need for more rigorous validation of detection tools before implementation in academic policy.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations