TemporalXAI-Det: Temporal-Aware Explainable Detection of Multi-Model AI-Generated Academic Text via Continual Learning and Cross-Lingual Transfer
Imeldawaty Gultom, Ratih Puspadini, Fauzi Erwis, Elyandri Prasiwiningrum, Ridwan · JOURNAL OF ICT APLICATIONS AND SYSTEM · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.56313/jictas.v5i1.531
Methodology & findings
Study design
Empirical evaluation study using a multi-stage framework with six integrated modules: (1) stylometric feature extraction, (2) cross-lingual semantic encoding, (3) hybrid feature fusion, (4) deep learning classification, (5) continual learning adaptation, and (6) explainability suite.
Sample
N = 72000, 7 groups
Primary method
McNemar's test (α = 0.05) with Bonferroni correction for pairwise statistical significance assessment. Stratified sampling for train/validation/test partitioning. AdamW optimizer with cosine annealing for model training (initial learning rate 3×10⁻⁵, batch size 64, 60 epochs, early stopping patience 10). Focal loss (γ = 2) for class-level imbalance handling within adversarial subsets. k-means clustering for experience replay buffer selection. Fisher information matrix for Elastic Weight Consolidation (EWC) regularization coefficient λ = 4,000.
Main result
TemporalXAI-Det achieves "97.2% accuracy and a macro F1 of 0.941" on clean test sets and demonstrates "a 78.4% reduction in forgetting achieved by EWC+Replay relative to standard fine-tuning" in continual learning scenarios. Under adversarial attacks, the system "exhibits a mean performance degradation of Δ = 2.9 pp across all attack conditions, compared to a mean of 24.6 pp across baselines," and the "LAPT mechanism achieves a mean adversarial macro F1 of 0.887 across eleven non-English languages using only 5% language-specific parameters."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist (computational experiments with quantitative measurement)
Author conclusions
The authors conclude that "meaningful LLM-family attribution is achievable at high accuracy (97.2% clean, 94.1% adversarial) even under sophisticated paraphrasing attacks" and that "the continual learning results demonstrate, for the first time in the AI-text detection literature, that catastrophic forgetting represents a quantifiable and addressable threat to deployed detection systems." They further note that "the LAPT mechanism achieves a mean adversarial macro F1 of 0.887 across eleven non-English languages using only 5% language-specific parameters," carrying "substantial equity implications" for global deployment.
Risk of bias
Potential selection bias in human-authored text sources (ArXiv, SSRN, PubMed, ERIC) which may not represent all academic disciplines equally; model selection bias toward five major LLM families; potential translation bias in multilingual extension; interannotator agreement at κ = 0.84 suggests moderate but not perfect annotation reliability.; Potential selection bias: AI text generation controlled via structured prompts derived from human texts; Temporal bias: Evaluation simulated temporal drift using predetermined phases rather than actual real-world model releases; Commercial detector evaluation: GPTZero and Turnitin APIs evaluated in March 2026; performance may reflect API-specific configurations; Class imbalance handling: Focal loss used to address imbalance within adversarial subsets, but equal class sizes may not reflect real-world distribution
Open questions raised
- The authors identify that binary human-vs-AI classifiers are inadequate for multi-source detection
- temporal model drift has been neglected by existing literature
- multilingual detection infrastructure is absent for non-English languages
- and future work on LLM attribution should focus on sub-lexical stylometric profiling rather than purely semantic approaches.
- Authors identify three compounding inadequacies in prior work:
- model multiplicity—binary classifiers fail across multiple LLM families with cross-model generalization rates below 60%
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations