Are AI Detectors Good Enough? A Survey on Quality of Datasets With Machine-Generated Texts
German Gritsai, Anastasia Voznyuk, Andrey Grabovoy, Yu. V. Chekhovich · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2410.14677
Methodology & findings
Study design
Systematic scoping review of datasets from AI-generated text detection competitions and research papers, combined with computational experiments.
Main result
The study found that "all analysed datasets fail in one or another of our methods and do not allow to reliably estimate AI detectors." Specifically, while some methods achieve F1-scores close to 1.0 (e.g., DeBERTa on HC3 dataset with 0.998), other detectors on the same datasets show dramatically lower performance (e.g., DetectGPT on DAGPap22 with 0.333 F1-score), suggesting that high benchmark scores may come from poor dataset quality rather than detector reliability. The paper demonstrates through topological analysis that "texts from SemEval, PAN24 and MGT-1" show similar PHD distributions between human and machine-generated texts, indicating higher quality data, while others exhibit clear separation patterns indicating dataset bias.
Research paradigm
Empirical-computational
Author conclusions
The authors conclude: "In the current research, we discussed the problem of quality of datasets with AI-generated texts used for testing corresponding detectors. This problem is relevant, as the quality of test data directly influences the quality of widely used detectors." They further state: "Our contribution aims to facilitate a better understanding of the dynamics between human and machine text, which will ultimately support the integrity of information in an increasingly automated world." The authors emphasize that "We encourage researchers to propose their own ways for quality assessment, which will allow to create a comprehensive system of evaluation of the detection datasets."
Risk of bias
Selection bias in dataset choice - only English datasets and most widely-cited collections were selected; Measurement bias - different baseline methods were inaccessible for non-English texts, limiting generalizability; Model bias - DeBERTa showed notably higher scores than zero-shot methods (Binoculars, DetectGPT), potentially due to fine-tuning advantage; Text length bias - methods not suitable for short texts (acknowledged limitation affecting RuATD, TweepFake, AuTex23-es); Generator representation bias - datasets use different language models across different time periods, affecting comparability; Selection bias in dataset choice: Only datasets from major competitions and cited research papers were included, potentially omitting smaller or less-cited datasets; Language bias: Primarily English datasets analyzed; non-English datasets (Russian, Spanish, Multilingual) underrepresented; Generator bias: Datasets vary in which language models were used for generation, affecting generalizability; Temporal bias: Datasets span different years (2019-2025) with varying LLM capabilities; Domain bias: Different datasets focus on different text domains (scientific papers, essays, news, social media); Sampling bias: Only 1000 balanced test samples per dataset used for evaluation, potentially not representative of full dataset characteristics; Dataset composition bias: different datasets use different generative models and text domains; Selection bias in dataset curation: datasets come from different sources with varying quality standards; Model fine-tuning bias: mDeBERTa baseline was fine-tuned for each dataset while other detectors were not; Language bias: some detectors (Binoculars, DetectGPT) work only on English texts; Text length bias: methods fail on short texts, affecting datasets with shorter average lengths; Generator bias: older datasets use older models (GPT-2, GPT-3) while newer datasets use more recent models
Limitations
- "In our work we focused on the task of binary classification, thus suggested methods are not optimal for the task of detection of the hybrid AI-human content
- Also, some methods do not work properly on short texts, however, this is a known issue for short texts." Additionally, the paper notes that "KLTTS may not perform well with very short texts, since the internal method of computing PHD requires sufficiently long texts for stable computation," requiring authors to discard KLTTS scores for RuATD, AuTex23-es, and TweepFake datasets.
Open questions raised
- Need for comprehensive quality assessment methods for AI-generated text detection datasets
- Lack of robust evaluation frameworks that account for multiple structural features of text
- Gap between benchmark performance and real-world detection performance
- Limited understanding of why detectors achieve near-perfect scores on some datasets but fail on others
- Need for standardized dataset quality criteria across detection tasks
- Unexplored potential of using high-quality generated data to improve both detection models and training datasets themselves
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations