Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection
Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmark construction and empirical evaluation study.
Sample
N = 279794, 7 groups
Primary method
Empirical evaluation using accuracy and F1-AI metrics. The paper reports performance metrics across revision versions and includes delta calculations showing change from previous version. No formal statistical testing (e.g., significance tests) or confidence intervals are provided. Tokenizer-agnostic token-level annotations are aligned to subword vocabularies via character offsets using whitespace tokenization.
Main result
The study found that "AI-text detectability is not determined solely by the proportion of AI-edited content. Instead, detector performance varies substantially across edit operations, domains, and revision stages." Most significantly, "intermediate mixed-authorship versions can be more challenging than both purely human and heavily AI-edited endpoints, revealing non-monotonic detection behavior that is overlooked by existing static benchmarks."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "reliable AI-text detection requires moving beyond binary endpoint classification toward trajectory-aware and operation-aware evaluation. OpAI-Bench provides a controlled testbed for analyzing whether, when, and how AI-assisted writing becomes detectable under realistic progressive editing scenarios." They further state that "detector behavior depends strongly on edit operation, domain, generator, and detector family. In particular, mixed-authorship intermediate versions can be harder to detect than both human-written and heavily AI-edited endpoints, revealing non-monotonic failure modes that static evaluations miss."
Risk of bias
Selection bias: Only documents with at least 10 sentences retained, potentially excluding shorter-form documents; Generator bias: Limited to three in-distribution generators (GPT-5.4, GPT-5.4-nano, Gemini-2.5-Flash) for training, with only Qwen3-8B as held-out; Domain bias: Evaluation limited to four specific domains; generalization to other domains unclear; Ordering bias: Deterministic shuffle based on document ID, though authors claim this removes positional bias; Selection bias: Documents retained only if N(D) ≥ 10 sentences, potentially excluding shorter documents and affecting domain representation; Domain bias: Limited to four specific domains (essays, news, reports, abstracts), may not represent all writing types; Generator bias: Primarily uses GPT-5.4, GPT-5.4-nano, and Gemini-2.5-Flash; Qwen3-8B used only for evaluation, not training; Ordering bias: While a deterministic shuffle is used to remove positional bias, the fixed seed-based ordering may still interact with specific document structures; Operation-specific bias: Five operations selected may not represent all forms of AI-assisted editing; Domain selection bias: Only four domains included; other writing styles may not be represented; Deterministic sentence shuffling: Using document ID-seeded shuffling may introduce predictable patterns detectable by models; Generator-specific effects: Only three primary generators (GPT-5.4, Gemini-2.5-Flash, GPT-5.4-nano) used for training; Qwen3-8B used only for held-out evaluation; Document minimum threshold: Requiring ≥10 sentences excludes shorter documents that may have different properties; Manual verification burden: Automatic validation and manual verification of sentence outputs could introduce inconsistencies; Prompt engineering effects: Editing prompts designed for specific operations may bias generation characteristics
Limitations
- The paper acknowledges that "most benchmarks are constructed from static final outputs and do not preserve the intermediate revision states that lead to those outputs." Additionally, the study is limited to "four domains" which may not represent all writing contexts
- The paper notes that "intermediate mixed-authorship versions can be more challenging than both purely human and heavily AI-edited endpoints," suggesting detectors were not optimized for this setting
- The study also notes that "sentence-level supervision improves stability, but does not remove revision effects," indicating persistent limitations in detection across revision stages.
Open questions raised
- The authors identify the gap that "existing AI-text detection benchmarks largely focus on final outputs and provide limited understanding of how AI authorship signals emerge, accumulate, or disappear throughout the revision process." They note that prior benchmarks "do not preserve explicit intermediate versions, do not model cumulative version-to-version editing, and do not jointly control AI coverage, edit type, and multi-granularity provenance within a unified setup," and call for future detectors that "can localize and reason about partial, incremental, and operation-specific AI involvement."
- Need for trajectory-aware and operation-aware detection methods beyond binary endpoint classification
- Lack of detectors that can localize and reason about partial, incremental, and operation-specific AI involvement
- Limited understanding of how detectability evolves during progressive human-AI co-editing
- Need for detectors that remain stable across different edit operations and cumulative revision histories
- Most prior benchmarks focus on final outputs rather than intermediate revision trajectories
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations