12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark construction and empirical evaluation study.

Sample

N = 279794, 7 groups

Primary method

Empirical evaluation using accuracy and F1-AI metrics. The paper reports performance metrics across revision versions and includes delta calculations showing change from previous version. No formal statistical testing (e.g., significance tests) or confidence intervals are provided. Tokenizer-agnostic token-level annotations are aligned to subword vocabularies via character offsets using whitespace tokenization.

Main result

The study found that "AI-text detectability is not determined solely by the proportion of AI-edited content. Instead, detector performance varies substantially across edit operations, domains, and revision stages." Most significantly, "intermediate mixed-authorship versions can be more challenging than both purely human and heavily AI-edited endpoints, revealing non-monotonic detection behavior that is overlooked by existing static benchmarks."

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "reliable AI-text detection requires moving beyond binary endpoint classification toward trajectory-aware and operation-aware evaluation. OpAI-Bench provides a controlled testbed for analyzing whether, when, and how AI-assisted writing becomes detectable under realistic progressive editing scenarios." They further state that "detector behavior depends strongly on edit operation, domain, generator, and detector family. In particular, mixed-authorship intermediate versions can be harder to detect than both human-written and heavily AI-edited endpoints, revealing non-monotonic failure modes that static evaluations miss."

Risk of bias

Selection bias: Only documents with at least 10 sentences retained, potentially excluding shorter-form documents; Generator bias: Limited to three in-distribution generators (GPT-5.4, GPT-5.4-nano, Gemini-2.5-Flash) for training, with only Qwen3-8B as held-out; Domain bias: Evaluation limited to four specific domains; generalization to other domains unclear; Ordering bias: Deterministic shuffle based on document ID, though authors claim this removes positional bias; Selection bias: Documents retained only if N(D) ≥ 10 sentences, potentially excluding shorter documents and affecting domain representation; Domain bias: Limited to four specific domains (essays, news, reports, abstracts), may not represent all writing types; Generator bias: Primarily uses GPT-5.4, GPT-5.4-nano, and Gemini-2.5-Flash; Qwen3-8B used only for evaluation, not training; Ordering bias: While a deterministic shuffle is used to remove positional bias, the fixed seed-based ordering may still interact with specific document structures; Operation-specific bias: Five operations selected may not represent all forms of AI-assisted editing; Domain selection bias: Only four domains included; other writing styles may not be represented; Deterministic sentence shuffling: Using document ID-seeded shuffling may introduce predictable patterns detectable by models; Generator-specific effects: Only three primary generators (GPT-5.4, Gemini-2.5-Flash, GPT-5.4-nano) used for training; Qwen3-8B used only for held-out evaluation; Document minimum threshold: Requiring ≥10 sentences excludes shorter documents that may have different properties; Manual verification burden: Automatic validation and manual verification of sentence outputs could introduce inconsistencies; Prompt engineering effects: Editing prompts designed for specific operations may bias generation characteristics

Limitations

  • The paper acknowledges that "most benchmarks are constructed from static final outputs and do not preserve the intermediate revision states that lead to those outputs." Additionally, the study is limited to "four domains" which may not represent all writing contexts
  • The paper notes that "intermediate mixed-authorship versions can be more challenging than both purely human and heavily AI-edited endpoints," suggesting detectors were not optimized for this setting
  • The study also notes that "sentence-level supervision improves stability, but does not remove revision effects," indicating persistent limitations in detection across revision stages.

Open questions raised

  • The authors identify the gap that "existing AI-text detection benchmarks largely focus on final outputs and provide limited understanding of how AI authorship signals emerge, accumulate, or disappear throughout the revision process." They note that prior benchmarks "do not preserve explicit intermediate versions, do not model cumulative version-to-version editing, and do not jointly control AI coverage, edit type, and multi-granularity provenance within a unified setup," and call for future detectors that "can localize and reason about partial, incremental, and operation-specific AI involvement."
  • Need for trajectory-aware and operation-aware detection methods beyond binary endpoint classification
  • Lack of detectors that can localize and reason about partial, incremental, and operation-specific AI involvement
  • Limited understanding of how detectability evolves during progressive human-AI co-editing
  • Need for detectors that remain stable across different edit operations and cumulative revision histories
  • Most prior benchmarks focus on final outputs rather than intermediate revision trajectories
Data: OpAI-Bench benchmark (https://github.com/VILA-Lab/OpAI-Bench) - contains 279,794 versioned samples from 31,089 trajectories across 15,722 source documents; OpAI-BenchCode: https://github.com/VILA-Lab/OpAI-Bench; GitHub: https://github.com/VILA-Lab/OpAI-Bench; OpAI-BenchExtracted from: pdfAgreement 53%

Explore related topics

Related papers