AI-generated feedback on writing: insights into efficacy and ENL student preference
Juan Escalante, Austin Pack, Alex Barrett · International Journal of Educational Technology in Higher Education · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41239-023-00425-2
Methodology & findings
Study design
Two longitudinal studies: Study 1 used a six-week repeated measures quasi-experimental design with 48 ENL students (experimental group receiving GPT-4 feedback, control group receiving human tutor feedback).
Sample
N = 91, 7 groups
Primary method
SPSS 28 software was used. Analysis included: repeated-measure ANOVA (RM-ANOVA) with general linear model using pretest and posttest as within-subjects variables and group as between-subjects variable with Greenhouse-Giesser correction; Shapiro-Wilk tests for normality; Levene's test for equality of variances; independent samples t-tests for comparing group means at pretest and posttest; intraclass correlation coefficients (ICC) for inter-rater reliability; and thematic analysis of qualitative data by three independent researchers.
Main result
Results of study 1 showed no difference in learning outcomes between the two groups. "The between-subject variable of group did not have a significant effect on writing scores, suggesting that one method of feedback was not better than another in terms of scores." Study 2 results revealed a near even split in preference for AI-generated or human-generated feedback, with "the six-week average of the number of students that preferred feedback from human tutors was 18, which was slightly lower than the average of those that preferred AI-generated feedback (19.667)."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed methods (quantitative quasi-experimental + qualitative survey)
Author conclusions
"In light of the major findings highlighted above, we believe a mixed approach to providing feedback may be most beneficial for both language educators and students. By utilizing GenAI, language educators may be able to produce more detailed feedback in a shorter amount of time for each individual learner. Providing opportunities for students to discuss AI-generated feedback with a human tutor and ask follow up questions affords students with the benefits of each." The authors state: "The main implication of these studies is that the use of AI-generated feedback can likely be incorporated into ENL essay evaluation without affecting learning outcomes, although we recommend a blended approach that utilizes the strengths of both forms of feedback."
Risk of bias
Selection bias: Non-probability self-selection method used for recruitment; Hawthorne effect: Students aware they were participating in a study comparing AI and human feedback; Attrition: Varying questionnaire response rates across weeks (32-41 participants with average of 37.7); Institutional bias: Single institution setting may not represent diverse educational contexts; Tutor variability: Human tutors were 'paid trained English language tutors' but potential differences in tutoring quality not controlled; AI consistency: Single AI model (GPT-4) used; prompt engineering quality depends on manual prompt design; Attrition: Some variation in questionnaire responses across weeks (32-41 participants); Confounding: Different tutors in control group may have different feedback styles; Potential demand characteristics in study 2 where students knew they were being evaluated on preference; Single institution sample limits generalizability; Potential demand characteristics: Students aware they were receiving feedback from different sources; Attrition: Variable completion rates in Study 2 (32-41 responses per week, average 37.7); Confounding variables: Different tutors for control group vs. AI-generated feedback; lack of randomization; Testing effects: Repeated weekly writing assignments may lead to familiarity; Tutor variation: Control group received feedback from human tutors who may differ in quality and style
Limitations
- The authors acknowledge that "In this study, as students were emailed the feedback generated by the AI, the students had no opportunity to ask follow-up questions to the AI
- This is because we wanted to limit students' access to the AI and prevent potential misuses where students asked the AI to write their assignments for them." Additionally, the study was conducted at a single institution during a shortened semester, limiting generalizability
- The qualitative analysis did not yield insights into why preference means dipped in certain weeks, and the study did not examine long-term retention of writing improvements.
Open questions raised
- Whether ChatGPT and similar GenAI tools can effectively and reliably be used for automated writing evaluation purposes
- Whether learners will accept feedback from GenAI tools
- Whether prompt engineering practices can produce corrective feedback reliably from LLMs
- Whether students and teachers perceive feedback from ChatGPT as useful in educational settings
- How to incorporate AI-generated feedback with opportunities for follow-up questions and discussion with human tutors
- Effects of different prompt engineering approaches on AWE reliability
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations