Detección de Textos Generados por Inteligencia Artificial Utilizando Deep Learning
Dario Emilio Hernández Loza, Mauro Pantoja Gutiérrez, Jonathán de Jesús Estrella Ramírez, Juan Carlos Gómez Carranza · JÓVENES EN LA CIENCIA · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.15174/jc.2025.4848
Methodology & findings
Study design
Empirical study using transformer-based deep learning models evaluated through 5-fold stratified cross-validation.
Sample
N = 2174, 2 groups
Primary method
5-fold stratified cross-validation; median aggregation of performance metrics across folds; One Cycle Policy for learning rate scheduling; early stopping based on ROC-AUC metric; standard classification metrics computed from confusion matrices (accuracy, precision, recall, F1-score). No inferential statistical tests (e.g., t-tests, ANOVA) or software packages explicitly named.
Main result
The study found that "el modelo ALBERT alcanzó el mejor desempeño y rendimiento, requiriendo la menor cantidad de tiempo de entrenamiento y logrando valores de 0.988 en todas las métricas" (ALBERT model achieved the best performance and efficiency, requiring the least amount of training time and achieving values of 0.988 in all metrics). The ALBERT transformer-based model demonstrated superior capability for discriminating between AI-generated and human-written texts compared to other models tested.
Reports effect sizes.
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude that "la arquitectura de trasformadores presenta una buena capacidad de clasificación en este tipo de tarea" (the transformer architecture presents a good classification capacity for this type of task). The "principal contribución de este trabajo es la siguiente: un estudio sobre la validación de modelos de aprendizaje profundo basados en transformadores en la tarea de identificación de textos generados por IA" (principal contribution of this work is the following: a study on the validation of deep learning models based on transformers in the task of identifying AI-generated texts).
Risk of bias
Dataset imbalance mitigation: original dataset had 14,131 AI-generated texts vs. 1,087 human texts; authors randomly selected 1,087 AI texts to create balanced dataset, which may not represent real-world distribution; Language specificity: all documents in English; generalization to other languages unknown; Limited domain: texts sourced from news domain via PAN@CLEF 2024; may not generalize to other domains; Class imbalance bias: models show tendency to correctly classify AI-generated texts but make more errors on human texts; Selection bias: specific AI generation methods and news sources not fully documented; Selection bias: Dataset manually balanced by random selection of 1,087 AI-generated texts, potentially missing diversity in AI generation methods; Data composition bias: All documents in English only; limited to news domain from PAN@CLEF 2024 competition; Class representation bias: Models showed asymmetric performance, with tendency to classify human-written texts as AI-generated more frequently than vice versa (except ALBERT); Limited human text diversity: Only 1,087 unique human-written texts available, requiring random sampling of AI texts to balance dataset; Performance metric selection: Focus on classical metrics without exploration of other potentially informative evaluation approaches; Selection bias: Dataset created by random selection of 1,087 AI-generated documents from a larger pool of 14,131 to balance with 1,087 human texts, potentially not representative of all AI-generated content; Language bias: All texts are in English only; findings may not generalize to other languages; Domain bias: Texts limited to news domain from PAN@CLEF 2024 task; results may not generalize across domains; Imbalanced original data: Original dataset had 14,131 AI texts vs 1,087 human texts before balancing; Metric selection bias: Only classical metrics used; other metrics not explored
Limitations
- The authors state: "En el trabajo realizado se limita a un conjunto de 2174 textos, esto debido a la baja cantidad de textos humanos se tuvo que limitar a la poca diversidad de textos generados por Inteligencia artificial" (The work was limited to a set of 2,174 texts, and due to the low quantity of human texts, had to be limited by the low diversity of texts generated by artificial intelligence)
- Additional limitations include: "El análisis se centró en métricas clásicas (Exactitud, precisión, Exh, tiempos de entrenamiento y prueba y F1), quedando pendiente la exploración de otras métricas que podrían enriquecer la evaluación" (The analysis focused on classical metrics, leaving pending the exploration of other metrics that could enrich the evaluation) and "Los resultados obtenidos pueden no generalizarse a textos en otros idiomas, dado que los datos empleado es específico del idioma inglés" (The results may not generalize to texts in other languages, since the data used is specific to the English language).
Open questions raised
- The authors identify the following gaps: (1) exploration of other evaluation metrics beyond classical metrics (Accuracy, Precision, Recall, F1); (2) evaluation on texts in languages other than English; (3) testing on larger and more diverse datasets; (4) investigation of robustness across different AI generation models and domains; (5) analysis of why models tend to misclassify human texts as AI-generated.
- The authors identify several pending research directions: exploration of metrics beyond classical accuracy measures (precision, recall, F1); generalization to texts in languages other than English; investigation of why models show asymmetric performance patterns, particularly the tendency to misclassify human-written texts as AI-generated; and potential development of more robust evaluation methods for transformer-based AI detection models.
- Exploration of metrics beyond classical ones (accuracy, precision, recall, F1)
- Evaluation on texts in languages other than English
- Testing on datasets larger than 2,174 documents with greater diversity
- Analysis beyond the news domain
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations