Process-Oriented Evaluation of AI-Assisted Scientific Writing
Paolo Silva, Sanchaita Hazra, D Lee, S Naresh Kumar, Bodhisattwa Prasad Majumder · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale keystroke-level log analysis with behavioral annotation.
Sample
N = 869, 8 groups
Primary method
Linear models predicting standardized change in linguistic properties with fixed effects for abstract provenance and pre-edit score, controlling for burst position and source paper. Gaussian mixture modeling (16 clusters) for burst behavior classification using Bayesian Information Criterion (BIC). Multinomial logistic regression examining education-moderated effects on burst behavior families. Paired t-tests (parametric) for pre-post linguistic comparisons between human and AI abstracts. Welch's t-tests for unequal variance comparisons between LM editors. Isometric log-ratio transformation for compositional features (action shares) prior to clustering. Effect sizes reported as Cohen's dz for paired designs. Multiple comparison correction applied (FDR).
Main result
The study found that "AI abstracts exhibit higher sentence-level agency, whereas human-authored abstracts outperform in global coherence, even with edits." Additionally, "human editors uphold local agency (p < .01) and structure (p < .01) in AI abstracts, but global coherence falls behind (p < .01) in human abstracts." The research demonstrates that "Language Models (LMs) improve edit outcomes through a mix of local and global features, but still actively struggle with global coherence."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical-quantitative with process-oriented behavioral analysis
Author conclusions
The authors conclude that "LMs are reshaping writing and editing, but little is known about how authors revise AI-generated scientific text." They emphasize: "we push the community towards a process revision mindset, using bursts in edit trajectories to reveal patterns in editing strategy. We highlight heterogeneous editing patterns by human editors at different scopes (word to sentence), but it is often ineffective to change an AI-generated text extensively, suggesting broader implications of behavior homogenization. While LMs tend to improve sentence-level features during editing, the benefits remain concentrated on the lowest-scoring abstracts, telling a cautionary tale of LM assistance in scientific writing."
Risk of bias
Selection bias: Participants were domain experts self-selected from author and reviewer pools with specific incentives; Disclosure bias: Authors explicitly varied disclosure of abstract source (AI vs. human), introducing demand characteristics; Education confounding: Education level moderated behavioral responses to AI disclosure, potentially confounding source effects; Attrition: 22 editing sessions deemed AI-assisted were dropped, reducing sample from 891 to 869; Measurement bias: Keystroke logs capture only digital editing; offline review/planning is unmeasured; Selection bias: Participants self-selected or were assigned from a limited pool of domain experts in Computer Science; generalizability to other disciplines unknown.; Confounding by education: Masters+ degree holders showed dramatically different behavior patterns; education effects on editing behavior not fully controlled.; Disclosure stigma: The 2×2 design explicitly manipulates whether participants know the abstract source is AI; this may induce demand characteristics or performative editing rather than authentic revision behavior.; Measurement reactivity: Keystroke logging itself may alter natural editing behavior; participants aware of monitoring may edit more carefully or strategically.; Attrition: 22 editing sessions dropped because they were 'deemed AI-assisted,' creating selection bias in the retained sample.; LM comparison validity: GPT 5.4 and Claude Opus 4.6 are future models (paper dated 2026); reproducibility is time-dependent.; Selection bias: Participants were domain experts with relevant experience; results may not generalize to non-experts or student writers.; Treatment effects: Participants knew they were being studied (keystroke logging), which may affect natural editing behavior.; Disclosure bias: The 2×2 between-subject design with disclosure treatment may prime stigmatic behavior against AI text.; Clustering stability: Manual labeling of 16 clusters introduces subjective interpretation risk; alternative clustering solutions not reported.; Confounding: Education level moderates behavior, but not controlled for in all analyses; demographic differences may correlate with unmeasured factors.
Open questions raised
- Limited understanding of how authors revise AI-generated scientific text (revision and refinement)
- Lack of process-oriented evaluation linking final text quality to revision behavior sequences
- Gap between final outputs and local revision behavior without explicit modeling of how draft properties drive edits
- Need for understanding of LM alignment with human editing intentions at different scopes
- Insufficient knowledge of macro-level behavior homogenization effects from repeated exposure to AI-generated text
- The authors identify that prior work has focused on final judgments of quality or isolated edit operations, but "two lingering questions remain. First, the existing models evaluate either final text quality or local revision behavior, without explicitly relating the two. Second, the revisions are often represented as isolated edit operations or aligned draft pairs. This obscures how draft-level properties drive subsequent edits and how sequences of edits alter a document over time." The paper also suggests future work should focus on "aligning LMs to faithfully follow human editing intentions" and understanding the concentration of stigma against AI text among higher-education groups.
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Practical and ethical challenges of large language models in education: A systematic scoping reviewLixiang Yan · 2023 · 699 citations