TemporalXAI-Det: Temporal-Aware Explainable Detection of Multi-Model AI-Generated Academic Text via Continual Learning and Cross-Lingual Transfer
Imeldawaty Gultom, Ratih Puspadini, Fauzi Erwis, Elyandri Prasiwiningrum, Ridwan · JOURNAL OF ICT APLICATIONS AND SYSTEM · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.56313/jictas.v5i1.531
Methodology & findings
Study design
Empirical evaluation study using a multi-stage framework with six integrated modules: (1) stylometric feature extraction, (2) cross-lingual semantic encoding, (3) hybrid feature fusion, (4) deep learning classification, (5) continual learning adaptation, and (6) explainability suite.
Sample
N = 72000, 16 groups
Primary method
McNemar's test (α = 0.05) with Bonferroni correction for pairwise statistical significance assessment. Stratified sampling for train/validation/test partitioning. AdamW optimizer with cosine annealing for model training (initial learning rate 3×10⁻⁵, batch size 64, 60 epochs, early stopping patience 10). Focal loss (γ = 2) for class-level imbalance handling within adversarial subsets. k-means clustering for experience replay buffer selection. Fisher information matrix for Elastic Weight Consolidation (EWC) regularization coefficient λ = 4,000.
Main result
TemporalXAI-Det achieves "97.2% accuracy and a macro F1 of 0.941" on clean test sets and demonstrates "a 78.4% reduction in forgetting achieved by EWC+Replay relative to standard fine-tuning" in continual learning scenarios. Under adversarial attacks, the system "exhibits a mean performance degradation of Δ = 2.9 pp across all attack conditions, compared to a mean of 24.6 pp across baselines," and the "LAPT mechanism achieves a mean adversarial macro F1 of 0.887 across eleven non-English languages using only 5% language-specific parameters."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist (computational experiments with quantitative measurement)
Author conclusions
The authors conclude that "meaningful LLM-family attribution is achievable at high accuracy (97.2% clean, 94.1% adversarial) even under sophisticated paraphrasing attacks" and that "the continual learning results demonstrate, for the first time in the AI-text detection literature, that catastrophic forgetting represents a quantifiable and addressable threat to deployed detection systems." They further note that "the LAPT mechanism achieves a mean adversarial macro F1 of 0.887 across eleven non-English languages using only 5% language-specific parameters," carrying "substantial equity implications" for global deployment.
Risk of bias
Potential selection bias in human-authored text sources (ArXiv, SSRN, PubMed, ERIC) which may not represent all academic disciplines equally; model selection bias toward five major LLM families; potential translation bias in multilingual extension; interannotator agreement at κ = 0.84 suggests moderate but not perfect annotation reliability.; Potential selection bias: AI text generation controlled via structured prompts derived from human texts; semantic fidelity validation (BERTScore ≥ 0.88) may introduce bias toward semantically similar paraphrases. Annotator bias in human verification (5% of samples, κ = 0.84). No discussion of potential dataset bias toward STEM and academic domains. Translation validation may not capture all linguistic nuances.; Selection bias: Human-authored texts drawn from four specific databases (arXiv, SSRN, PubMed, ERIC) may not represent all academic writing patterns; AI text generation bias: All AI texts generated from prompts derived from human-authored abstracts, potentially biasing generators toward human-like outputs; Temporal bias: Evaluation simulated temporal drift using predetermined phases rather than actual real-world model releases; Multilingual translation bias: All non-English samples generated via machine translation (NLLB-200) rather than native generation in target languages; Annotation bias: Interannotator agreement κ = 0.84 for 5% sample verification indicates moderate but not perfect agreement; Commercial detector evaluation: GPTZero and Turnitin APIs evaluated in March 2026; performance may reflect API-specific configurations; Class imbalance handling: Focal loss used to address imbalance within adversarial subsets, but equal class sizes may not reflect real-world distribution
Open questions raised
- The authors identify that binary human-vs-AI classifiers are inadequate for multi-source detection; temporal model drift has been neglected by existing literature; multilingual detection infrastructure is absent for non-English languages; and future work on LLM attribution should focus on sub-lexical stylometric profiling rather than purely semantic approaches.
- Authors identify three compounding inadequacies in prior work: (1) model multiplicity—binary classifiers fail across multiple LLM families with cross-model generalization rates below 60%; (2) temporal drift—detectors trained in Q1 2024 show 15-22 percentage point accuracy degradation against models released in Q4 2024; (3) linguistic coverage—existing detectors overwhelmingly focus on English, creating equity gaps for institutions in the Global South. Future work should "focus on sub-lexical stylometric profiling rather than purely semantic approaches" for LLM attribution.
- Future work on LLM attribution should focus on sub-lexical stylometric profiling rather than purely semantic approaches
- No prior work has applied continual learning specifically to temporal model drift in AI-text detection
- Cross-lingual transfer for AI-text detection is nascent; prior work (Liao et al.) showed RoBERTa-based detectors fine-tuned on English achieve near-random performance on Mandarin AI text
- Existing AI-text detectors are overwhelmingly designed and evaluated on English text, creating a structural equity gap for Global South institutions
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations