12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Process-Oriented Evaluation of AI-Assisted Scientific Writing

Paolo Silva, Sanchaita Hazra, D Lee, S Naresh Kumar, Bodhisattwa Prasad Majumder · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale keystroke-level log analysis with behavioral annotation.

Sample

N = 869, 8 groups

Primary method

Linear models predicting standardized change in linguistic properties with fixed effects for abstract provenance and pre-edit score, controlling for burst position and source paper. Gaussian mixture modeling (16 clusters) for burst behavior classification using Bayesian Information Criterion (BIC). Multinomial logistic regression examining education-moderated effects on burst behavior families. Paired t-tests (parametric) for pre-post linguistic comparisons between human and AI abstracts. Welch's t-tests for unequal variance comparisons between LM editors. Isometric log-ratio transformation for compositional features (action shares) prior to clustering. Effect sizes reported as Cohen's dz for paired designs. Multiple comparison correction applied (FDR).

Main result

The study found that "AI abstracts exhibit higher sentence-level agency, whereas human-authored abstracts outperform in global coherence, even with edits." Additionally, "human editors uphold local agency (p < .01) and structure (p < .01) in AI abstracts, but global coherence falls behind (p < .01) in human abstracts." The research demonstrates that "Language Models (LMs) improve edit outcomes through a mix of local and global features, but still actively struggle with global coherence."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical-quantitative with process-oriented behavioral analysis

Author conclusions

The authors conclude that "LMs are reshaping writing and editing, but little is known about how authors revise AI-generated scientific text." They emphasize: "we push the community towards a process revision mindset, using bursts in edit trajectories to reveal patterns in editing strategy. We highlight heterogeneous editing patterns by human editors at different scopes (word to sentence), but it is often ineffective to change an AI-generated text extensively, suggesting broader implications of behavior homogenization. While LMs tend to improve sentence-level features during editing, the benefits remain concentrated on the lowest-scoring abstracts, telling a cautionary tale of LM assistance in scientific writing."

Risk of bias

Selection bias: Participants were domain experts self-selected from author and reviewer pools with specific incentives; Disclosure bias: Authors explicitly varied disclosure of abstract source (AI vs. human), introducing demand characteristics; Education confounding: Education level moderated behavioral responses to AI disclosure, potentially confounding source effects; Attrition: 22 editing sessions deemed AI-assisted were dropped, reducing sample from 891 to 869; Measurement bias: Keystroke logs capture only digital editing; offline review/planning is unmeasured; Selection bias: Participants self-selected or were assigned from a limited pool of domain experts in Computer Science; generalizability to other disciplines unknown.; Confounding by education: Masters+ degree holders showed dramatically different behavior patterns; education effects on editing behavior not fully controlled.; Disclosure stigma: The 2×2 design explicitly manipulates whether participants know the abstract source is AI; this may induce demand characteristics or performative editing rather than authentic revision behavior.; Measurement reactivity: Keystroke logging itself may alter natural editing behavior; participants aware of monitoring may edit more carefully or strategically.; Attrition: 22 editing sessions dropped because they were 'deemed AI-assisted,' creating selection bias in the retained sample.; LM comparison validity: GPT 5.4 and Claude Opus 4.6 are future models (paper dated 2026); reproducibility is time-dependent.; Selection bias: Participants were domain experts with relevant experience; results may not generalize to non-experts or student writers.; Treatment effects: Participants knew they were being studied (keystroke logging), which may affect natural editing behavior.; Disclosure bias: The 2×2 between-subject design with disclosure treatment may prime stigmatic behavior against AI text.; Clustering stability: Manual labeling of 16 clusters introduces subjective interpretation risk; alternative clustering solutions not reported.; Confounding: Education level moderates behavior, but not controlled for in all analyses; demographic differences may correlate with unmeasured factors.

Open questions raised

  • Limited understanding of how authors revise AI-generated scientific text (revision and refinement)
  • Lack of process-oriented evaluation linking final text quality to revision behavior sequences
  • Gap between final outputs and local revision behavior without explicit modeling of how draft properties drive edits
  • Need for understanding of LM alignment with human editing intentions at different scopes
  • Insufficient knowledge of macro-level behavior homogenization effects from repeated exposure to AI-generated text
  • The authors identify that prior work has focused on final judgments of quality or isolated edit operations, but "two lingering questions remain. First, the existing models evaluate either final text quality or local revision behavior, without explicitly relating the two. Second, the revisions are often represented as isolated edit operations or aligned draft pairs. This obscures how draft-level properties drive subsequent edits and how sequences of edits alter a document over time." The paper also suggests future work should focus on "aligning LMs to faithfully follow human editing intentions" and understanding the concentration of stigma against AI text among higher-education groups.
Data: The paper states "Code, Data" are available but does not provide explicit URLs or repository links in the excerpt provided.; The paper states "Code, Data" are available but does not provide explicit URLs or repositories. Data sourced from Hazra et al. (2026).; The paper states "Code, Data" are available but does not provide explicit URLs. The paper references data from Hazra et al. (2026), which is the source corpus.Code: The paper references "Code, Data" in the header but does not provide specific GitHub or GitLab repository URLs in the document.; Not explicitly provided.; The paper mentions "Code, Data" are available but does not provide GitHub or GitLab repository links in the accessible text.Extracted from: pdfAgreement 39%

Explore related topics

Related papers