Co-Writing with AI: An Empirical Study of Diverse Academic Writing Workflows
Silvia Bodei, Duncan P. Brumby, Katie Fisher, Jon Mella · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3808045.3808072
Methodology & findings
Study design
Mixed methods study combining (1) quantitative online survey with 107 UK university students (undergraduates n=52, master's students n=48, doctoral students n=7) measuring stage-specific AI use and individual factors using Likert scales and checklists; (2) semi-structured interviews with 12 postgraduate participants from an HCI program with explicit AI training, using visual probes and open-ended exploration of writing workflows..
Sample
N = 119, 7 groups
Primary method
Analyses conducted in R. Methods included: (1) Descriptive statistics summarizing background measures, AI literacy, writing confidence, and stage-specific usage proportions; (2) Welch's independent-samples t-tests (chosen for robustness to violations of normality and homogeneity of variance in smaller samples) comparing mean scores between AI users and non-users at each writing stage; (3) Pairwise deletion approach for missing data under Missing at Random (MAR) assumption; (4) Qualitative thematic analysis of interview data organized around three workflow profiles (quality-oriented, learning-oriented, productivity-oriented).
Main result
The study found that "students generally concentrate AI use on a limited subset of the writing process, rather than uniformly applying it across all stages." More specifically, "AI use was relatively evenly distributed across writing stages, with no single stage emerging as the dominant point of reliance," and the "average number of stages with reported AI use was M = 2.38, SD = 1.26." The research identified three distinct workflow configurations: early-stage use centered on ideation, sourcing, and planning; late-stage use centered on drafting and reviewing; and cross-stage use linking early and later stages. Importantly, "what appeared to distinguish AI users from non-users was not writing confidence but how they evaluated the tool's benefits, risks, and legitimacy."
Reports effect sizes.
Research paradigm
Mixed methods (quantitative survey + qualitative interviews)
Author conclusions
"This research shows that AI integration in academic writing is neither uniform nor reducible to binary notions of 'critical' and 'non-critical' use. Instead, students organize AI use in stage-specific ways, selectively integrating it across writing workflows in response to competing priorities such as learning, quality, productivity, and authorship." The authors conclude that "Rather than enforcing uniform models of use, future frameworks, tools, and policies should recognize this variability and focus on supporting informed, context-sensitive integration. As AI makes text production easier, the central challenge becomes not whether students can generate writing, but how they engage with, evaluate, and take responsibility for it."
Risk of bias
Selection bias: Convenience sampling used in both studies; participants from UK institutions only; self-selection effects (those willing to discuss AI use); Social desirability bias: Self-reported AI use may underestimate actual behavior; students may conceal use due to reputational risks; Attrition bias: 186 participants initiated survey; only 107 completed ≥80% (42% exclusion rate); Disclosure bias: Moderate comfort with disclosing AI use (M=3.15, SD=1.20); lowest comfort disclosing to lecturers (M=2.77); Sample representativeness: Interview sample (n=12) from single institution with explicit AI support; high AI literacy and writing confidence ceiling effects; Missing data: Available-case approach used for factor analyses due to survey attrition; Researcher positionality: Interviewer peer status may have reduced evaluative pressure but shaped what participants felt appropriate to disclose; Selection bias: convenience sampling via Prolific and personal networks; interview participants from supportive institutional context with explicit AI literacy training; Social desirability bias: students may under-report AI use due to policy ambiguity and social dynamics; Disclosure bias: self-reported AI use shaped by trust, credibility, and researcher-participant relationship dynamics; Attrition bias: 79 of 186 survey initiates excluded (42.5% incomplete responses); Sampling bias: Study 2 participants exhibited ceiling effects for AI literacy and writing confidence; Confounding: disclosure comfort, departmental AI policies, prior education not controlled statistically; Selection bias: Convenience sampling in both studies; Social desirability bias: Self-reported AI use may underestimate actual behavior; Disclosure bias: Participants may conceal AI use due to policy ambiguity or perceived risk; Attrition bias: Survey excluded participants with <80% completion (n=79 excluded from 186); Sample representativeness: Study 2 sample consisted of students with explicit AI training and supportive institutional context, limiting generalizability; Ceiling effects: Interview participants exhibited ceiling effects for AI literacy and writing confidence relative to survey sample
Limitations
- "Although appropriate for an exploratory study, the reliance on single ad hoc items limits the robustness and interpretability of the measured constructs
- In addition, sample size and power constraints restrict the analysis of more complex relationships between factors." Additionally, "self-report measures provide limited control over how participants interpret constructs such as 'use,' 'trust,' or 'authorship,' and may not capture the contextual reasoning underlying their responses
- These challenges are particularly salient in the context of AI use in education, where reported use may underestimate actual behavior due to social desirability bias and selective disclosure." The interview sample was "limited in size and composed of students with high AI literacy and writing confidence, which restricts the generalisability of the workflow profiles identified," and "the profiles identified here may therefore over-represent deliberate, well-regulated forms of integration."
Open questions raised
- Limited research explicitly connecting individual-level factors to where within the writing process AI is used
- Few studies examining how students organize AI use across complete writing workflows and stages
- Lack of understanding of how contextual and individual variability shapes stage-specific AI integration as AI becomes more embedded in practice
- Need for comparative evidence on stage-based patterns across more diverse populations, disciplines, and institutional contexts
- Limited research using validated multi-item scales and larger samples to support robust modeling of relationships between individual factors and AI use patterns
- Need for longitudinal and naturalistic approaches to understand how decisions about delegation evolve over time
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations