12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

From digital divide to equity-enhancing diffusion: Generative AI and writing quality

Rebecca Tukachinsky Forster, Kerk F. Kee, Gabriel Miao Li · AI & Society · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
E
Evidence
2
Citations
0.87
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02739-3

Methodology & findings

Study design

Within-subject experimental design where students (N=170) wrote two short essays (100 words each), one with and one without ChatGPT assistance.

Sample

N = 170, 6 groups

Primary method

Paired t-tests comparing AI-assisted versus non-assisted essays; chi-square tests comparing dichotomous variables (multiple prompts, prompt personalization) between strong and developing writers; independent samples t-tests for output modification comparisons; linear mixed-effects regression models (lme4 package in R, p-values via lmerTest) accounting for cross-classified random factors (participant and topic); mediation analyses using structural equation modeling (lavaan package in R); logistic and ordinary least squares regressions in sensitivity analyses. Cronbach's alpha for internal consistency; Krippendroff's alpha for inter-rater reliability.

Main result

The study found that "while all students benefited from AI, that less skillful writers gained more." Specifically, developing writers gained on average 9.03 points (SD = 10.43) from using ChatGPT compared to stronger writers who gained only 3.27 points (SD = 8.36), with a significant difference (t(168) = 3.88, p < 0.001, Cohen's d = 0.60). For human assessment, "strong writers only improved their score by an average of 1.72 points (SD = 7.79) when using AI, but the developing writers gained on average 13.93 (SD = 11.18) points."

Reports effect sizes and confidence intervals.

Research paradigm

post-positivist empiricism with quantitative and qualitative measurement

Author conclusions

The authors conclude: "Contributing to the literature on digital divide and knowledge gap in diffusion research, this study demonstrated the potential of generative AI to level the playing field for skilled and developing business communication writers. ChatGPT assisted all students, but it was particularly helpful for students who were not strong writers, improving their scores by almost a full grade." They further argue that "innovation consequences are not fixed properties of technologies but interactional achievements emerging from communicative implementation" and that "generative AI can enhance both efficiency and organizational equity."

Risk of bias

Selection bias: Participants self-selected into writing center and communication courses; only US citizens and permanent residents eligible for monetary compensation; Attrition: 170 final participants after removing incomplete data and non-compliant participants; Rater bias: Human coders were manuscript authors rather than fully independent raters, though 10% of sample double-coded for reliability; Order effects: Small but significant order effect found for computerized scores (B = 1.597, SE = 0.75, p = 0.034); Generalizability: Limited to short business communication writing task; results may not generalize to longer or creative writing; Selection bias: Recruitment from communication courses and writing centers may not represent general student population; Restriction of sample: Only US citizens and permanent residents eligible for monetary compensation; Coder bias: Human coders were authors of the manuscript, though blind to condition and training was conducted; Order effect: Small but significant order effect detected for computerized scores (B = 1.597, p = 0.034); Attrition: Incomplete data and non-compliance resulted in removal of participants before final sample of 170; Selection bias: Participants self-selected from introductory communication courses and writing centers, potentially skewing toward more academically engaged students; Demographic restriction: Only US citizens and permanent residents were eligible for monetary compensation due to university and immigration regulations; Coder bias: Human coding was conducted by two of the manuscript authors rather than independent external raters, though coders were blinded to condition; Order effect: Small but significant order effect detected in computerized scores (B = 1.597, SE = 0.75, p = 0.034), suggesting counterbalancing may not have fully eliminated sequential effects; Hawthorne effect: Promise of cash prize for top 5% essays may have motivated participants differently; Limited generalizability: Findings specific to 100-word business communication essays; unclear if effects generalize to other writing tasks or genres

Limitations

  • The study "was specifically focused on writing short texts with a specific genre (i.e., business communication)
  • It is possible that ChatGPT would be less helpful with longer writing tasks or tasks that involve different types of prompts and/or genres (e.g., creative, technical, persuasive)." Additionally, "the study is limited to using ChatGPT5" and "was limited to a general assessment of writing quality using a relatively crude overall Grammarly score and various human rating scales that were implemented by the manuscript authors." Furthermore, "this study was not designed to empirically test the concerns that have been raised regarding possible long-term risks of over-reliance on AI (e.g., Bai et al

Open questions raised

  • Generalizability to longer writing tasks and different genres (creative, technical, persuasive)
  • Testing with other generative AI platforms (Claude, later ChatGPT versions)
  • More nuanced understanding of specific dimensions of writing most enhanced by AI
  • Longitudinal studies examining whether performance gains constitute durable skill development or short-term tool-dependent improvements
  • Empirical testing of concerns regarding long-term risks of AI over-reliance
  • Broader research agenda examining how human-AI co-production can foster equitable participation in knowledge economies
Data: Study data; Data available at: https://doi.org/10.60911/chapman.30604586Code: Not mentionedExtracted from: pdfAgreement 52%

Explore related topics

Related papers