12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Word Embeddings and Stylistic Features based Approach for Generative AI Authorship Verification

T. S. S. R. K. Rao, K.Upendra Raju, Vivek Krishna, S. Chanti, Karunakar Kavuri · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1109/eaic66483.2025.11101533

Methodology & findings

Study design

Experimental machine learning approach using a hybrid model combining stylistic features and word embeddings from transformer models (RoBERTa and T5) with LightGBM classification algorithm, evaluated on the PAN 2024 Generative AI Authorship Verification dataset..

Sample

unknown, 2 groups

Primary method

LightGBM algorithm for training the hybrid classification model. Evaluation metrics include ROC-AUC, C@1, Brier score, F0.5u, F1, and mean.

Main result

The proposed hybrid model attained strong performance metrics on the training dataset: "The proposed hybrid model attained scores of 0.996 for ROC-AUC, 0.970 for Brier, 0.989 for C@1, 0.972 for F1, 0.971 for F0.5u, and 0.979 for mean on the training dataset." These results demonstrate the effectiveness of combining transformer embeddings with stylistic features for authorship verification tasks.

Reports effect sizes.

Research paradigm

positivist/empiricist

Author conclusions

The authors conclude that "the proposed hybrid model attained scores of 0.996 for ROC-AUC, 0.970 for Brier, 0.989 for C@1, 0.972 for F1, 0.971 for F0.5u, and 0.979 for mean on the training dataset. These results demonstrate the importance of including transformer embeddings along with stylistic features to improve the performance of authorship verification."

Risk of bias

No mention of cross-validation or test set performance - only training dataset results reported; No discussion of class imbalance in the dataset; Potential overfitting risk given very high training metrics (0.996 ROC-AUC); No comparison with baseline models or prior work mentioned; Dataset composition and sources not fully detailed; Potential overfitting: evaluation metrics reported only on training data, not on held-out test or validation sets; Unclear data sampling methodology and potential selection bias in dataset composition; No discussion of class imbalance or demographic representation of different LLM sources; Limited transparency on hyperparameter tuning procedures

Data: not_statedCode: not_statedExtracted from: pdfAgreement 67%

Explore related topics

Related papers