Generative feedback: causal effects of LLM’s on writing quality and evaluative equity in higher education
Carelys Suescum Coelho, Car-Emyr Suescum Coelho, Carluys Suescum Coelho, Carlysmar Suescum Coelho · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.62486/978-9915-9854-5-9.ch06
Methodology & findings
Study design
This is a theoretical chapter proposing a recommended methodological design for future research.
Primary method
The chapter recommends: multilevel models (students at level 1, sections/classes at level 2), regressions with subgroup dummy variables and their interactions with treatment, conditional effect estimation (CATE) alongside average treatment effects (ATE), statistical correction methods for multiple comparisons (Bonferroni or FDR), and sensitivity analyses. Analysis of effect heterogeneity would examine whether treatment reduces score gaps between groups.
Main result
The study indicates that "generative feedback, when properly designed, can significantly improve the revision cycle and enhance writing quality" and that "by applying consistent criteria, AI can level the playing field for students with less linguistic proficiency or who are first-generation students, giving them access to the same constructive feedback as their more experienced peers, thereby reducing dispersion in results." The authors note "consistent gains when LLM feedback is effectively integrated with explicit rubrics and revision microtasks; it also increases engagement and can decrease inter-rater variability."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed-methods (positivist/empiricist for proposed RCT design; interpretivist for ethical/equity framework)
Author conclusions
The authors conclude that "LLM-powered generative feedback has considerable potential to improve writing quality in secondary and higher education by providing detailed, timely comments that leverage students' self-regulated learning" and that "this approach can contribute to greater evaluative equity by applying consistent criteria, thereby reducing baseline disparities between groups and expanding access to meaningful feedback. However, it is important to note that these positive effects depend on careful design that requires human oversight and transparent architecture to avoid automatic biases." They envision "a future in which writing excellence and equity are enhanced by responsible AI" through "iterative improvement paths that monitor new LLM developments, integrate automated feedback into curricula, and train the entire academic community in its critical use."
Risk of bias
Model version sensitivity - effectiveness depends on specific LLM version used at time of implementation; Algorithmic bias - potential biases in LLM responses, particularly favoring native speaker linguistic patterns; Rubric subjectivity - measurement of writing quality subject to rater interpretation despite blinding; Confounding by unobserved variables - L1/L2 status may be correlated with socioeconomic status; Contamination risk - control group may receive additional feedback outside of study protocol; Selection bias - generalizability limited by institutional context and technological infrastructure; Model version sensitivity and variability across LLM updates; Algorithmic biases in LLM-generated feedback; Subjectivity in writing quality rubric evaluation despite blinding; Unobserved confounding variables (e.g., socioeconomic status correlated with L1/L2 status); Potential reverse bias favoring native speakers' linguistic patterns; Contamination between control and treatment groups; Model version dependency - future LLM updates could alter feedback effectiveness; Algorithmic bias - LLMs may privilege linguistic patterns of native speakers; Rubric subjectivity - measurement of writing quality depends on analytical rubric reliability; Selection bias potential - generalizability limited by specific institutional contexts; Confounding variables - unobserved factors like socioeconomic status may correlate with L1/L2 categories; Contamination risk - control group may receive additional feedback
Limitations
- The authors state that "specific data and contexts, such as courses, subjects, institutions, and cohorts, may limit generalizability
- The results will depend on the LLM model version available during the experiment phase, and future updates to the model could alter the effectiveness of the feedback." Additionally, "the measurement of writing quality would depend on the reliability of the analytical rubrics
- Although the grading will be shielded, there will always be room for subjectivity." Furthermore, "in terms of equity, the categories analyzed for L1 and L2 student levels may be correlated with unobserved factors, such as socioeconomic status, which would complicate the interpretation of the gaps."
Open questions raised
- Lack of robust research on precise differential effects by subgroups
- Limited evidence on large-scale institutional deployment of generative feedback
- Need for longitudinal follow-ups to assess whether writing improvements persist over time
- Extension to other disciplines where writing has different functions
- Integration with broader learning analytics to enrich understanding of mechanisms of change
- Cost-benefit and institutional effectiveness evaluation for sustainable adoption
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations