A Word Embeddings and Stylistic Features based Approach for Generative AI Authorship Verification
T. S. S. R. K. Rao, K.Upendra Raju, Vivek Krishna, S. Chanti, Karunakar Kavuri · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1109/eaic66483.2025.11101533
Methodology & findings
Study design
Experimental machine learning approach using a hybrid model combining stylistic features and word embeddings from transformer models (RoBERTa and T5) with LightGBM classification algorithm, evaluated on the PAN 2024 Generative AI Authorship Verification dataset..
Sample
unknown, 2 groups
Primary method
LightGBM algorithm for training the hybrid classification model. Evaluation metrics include ROC-AUC, C@1, Brier score, F0.5u, F1, and mean.
Main result
The proposed hybrid model attained strong performance metrics on the training dataset: "The proposed hybrid model attained scores of 0.996 for ROC-AUC, 0.970 for Brier, 0.989 for C@1, 0.972 for F1, 0.971 for F0.5u, and 0.979 for mean on the training dataset." These results demonstrate the effectiveness of combining transformer embeddings with stylistic features for authorship verification tasks.
Reports effect sizes.
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "the proposed hybrid model attained scores of 0.996 for ROC-AUC, 0.970 for Brier, 0.989 for C@1, 0.972 for F1, 0.971 for F0.5u, and 0.979 for mean on the training dataset. These results demonstrate the importance of including transformer embeddings along with stylistic features to improve the performance of authorship verification."
Risk of bias
No mention of cross-validation or test set performance - only training dataset results reported; No discussion of class imbalance in the dataset; Potential overfitting risk given very high training metrics (0.996 ROC-AUC); No comparison with baseline models or prior work mentioned; Dataset composition and sources not fully detailed; Potential overfitting: evaluation metrics reported only on training data, not on held-out test or validation sets; Unclear data sampling methodology and potential selection bias in dataset composition; No discussion of class imbalance or demographic representation of different LLM sources; Limited transparency on hyperparameter tuning procedures
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations