From digital divide to equity-enhancing diffusion: Generative AI and writing quality
Rebecca Tukachinsky Forster, Kerk F. Kee, Gabriel Miao Li · AI & Society · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02739-3
Methodology & findings
Study design
Within-subject experimental design where students (N=170) wrote two short essays (100 words each), one with and one without ChatGPT assistance.
Sample
N = 170, 6 groups
Primary method
Paired t-tests comparing AI-assisted versus non-assisted essays; chi-square tests comparing dichotomous variables (multiple prompts, prompt personalization) between strong and developing writers; independent samples t-tests for output modification comparisons; linear mixed-effects regression models (lme4 package in R, p-values via lmerTest) accounting for cross-classified random factors (participant and topic); mediation analyses using structural equation modeling (lavaan package in R); logistic and ordinary least squares regressions in sensitivity analyses. Cronbach's alpha for internal consistency; Krippendroff's alpha for inter-rater reliability.
Main result
The study found that "while all students benefited from AI, that less skillful writers gained more." Specifically, developing writers gained on average 9.03 points (SD = 10.43) from using ChatGPT compared to stronger writers who gained only 3.27 points (SD = 8.36), with a significant difference (t(168) = 3.88, p < 0.001, Cohen's d = 0.60). For human assessment, "strong writers only improved their score by an average of 1.72 points (SD = 7.79) when using AI, but the developing writers gained on average 13.93 (SD = 11.18) points."
Reports effect sizes and confidence intervals.
Research paradigm
post-positivist empiricism with quantitative and qualitative measurement
Author conclusions
The authors conclude: "Contributing to the literature on digital divide and knowledge gap in diffusion research, this study demonstrated the potential of generative AI to level the playing field for skilled and developing business communication writers. ChatGPT assisted all students, but it was particularly helpful for students who were not strong writers, improving their scores by almost a full grade." They further argue that "innovation consequences are not fixed properties of technologies but interactional achievements emerging from communicative implementation" and that "generative AI can enhance both efficiency and organizational equity."
Risk of bias
Selection bias: Participants self-selected into writing center and communication courses; only US citizens and permanent residents eligible for monetary compensation; Attrition: 170 final participants after removing incomplete data and non-compliant participants; Rater bias: Human coders were manuscript authors rather than fully independent raters, though 10% of sample double-coded for reliability; Order effects: Small but significant order effect found for computerized scores (B = 1.597, SE = 0.75, p = 0.034); Generalizability: Limited to short business communication writing task; results may not generalize to longer or creative writing; Selection bias: Recruitment from communication courses and writing centers may not represent general student population; Restriction of sample: Only US citizens and permanent residents eligible for monetary compensation; Coder bias: Human coders were authors of the manuscript, though blind to condition and training was conducted; Order effect: Small but significant order effect detected for computerized scores (B = 1.597, p = 0.034); Attrition: Incomplete data and non-compliance resulted in removal of participants before final sample of 170; Selection bias: Participants self-selected from introductory communication courses and writing centers, potentially skewing toward more academically engaged students; Demographic restriction: Only US citizens and permanent residents were eligible for monetary compensation due to university and immigration regulations; Coder bias: Human coding was conducted by two of the manuscript authors rather than independent external raters, though coders were blinded to condition; Order effect: Small but significant order effect detected in computerized scores (B = 1.597, SE = 0.75, p = 0.034), suggesting counterbalancing may not have fully eliminated sequential effects; Hawthorne effect: Promise of cash prize for top 5% essays may have motivated participants differently; Limited generalizability: Findings specific to 100-word business communication essays; unclear if effects generalize to other writing tasks or genres
Limitations
- The study "was specifically focused on writing short texts with a specific genre (i.e., business communication)
- It is possible that ChatGPT would be less helpful with longer writing tasks or tasks that involve different types of prompts and/or genres (e.g., creative, technical, persuasive)." Additionally, "the study is limited to using ChatGPT5" and "was limited to a general assessment of writing quality using a relatively crude overall Grammarly score and various human rating scales that were implemented by the manuscript authors." Furthermore, "this study was not designed to empirically test the concerns that have been raised regarding possible long-term risks of over-reliance on AI (e.g., Bai et al
Open questions raised
- Generalizability to longer writing tasks and different genres (creative, technical, persuasive)
- Testing with other generative AI platforms (Claude, later ChatGPT versions)
- More nuanced understanding of specific dimensions of writing most enhanced by AI
- Longitudinal studies examining whether performance gains constitute durable skill development or short-term tool-dependent improvements
- Empirical testing of concerns regarding long-term risks of AI over-reliance
- Broader research agenda examining how human-AI co-production can foster equitable participation in knowledge economies
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations