Students’ Perceptions and Preferences of Generative Artificial Intelligence Feedback for Programming
Zhengdong Zhang, Zihan Dong, Yang Shi, Thomas Price, Noboru Matsuda, Dongkuan Xu · Proceedings of the AAAI Conference on Artificial Intelligence · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30372
Methodology & findings
Study design
Survey-based empirical study with thematic qualitative analysis.
Sample
N = 58, 5 groups
Primary method
Descriptive statistics (percentages, frequencies) calculated for Likert-scale survey responses. Thematic analysis conducted independently by three researchers with collaborative reconciliation of codes. No inferential statistical tests (e.g., t-tests, ANOVA, chi-square) reported. Percentages calculated for preference data and guideline alignment (e.g., 72.5%, 98%, 95%, 71%).
Main result
The study found that "the ChatGPT-generated feedback largely aligns with formative feedback guidelines, with over 70% of students giving favorable evaluations of this alignment." Additionally, "72.5% (74/102) of the responses preferred the feedback generated with the prompt that contains students' code (feedback 1) over the feedback without students' code (feedback 2)," with students citing specificity (33/74), clarity (15/74), and corrective suggestions (9/74) as the top three reasons for their preference.
Reports effect sizes.
Research paradigm
Empirical-interpretivist (mixed methods: survey + thematic analysis)
Author conclusions
"This study demonstrated that ChatGPT could generate Java programming assignment feedback that students perceived as formative. It also offered insights into the specific improvements that would make the ChatGPT-generated feedback useful for students." The authors note that "ChatGPT has considerable potential for generating practical and meaningful assignment feedback" and recommend that "educators could explore personalization options to better meet students' diverse needs in real-world educational contexts," as "student preferences for feedback tone varied. Some students favored a positive tone, while others sought more critical assessments."
Risk of bias
Selection bias: All participants were from a single introductory Java course at one U.S. public university, predominantly computer science majors; Attrition: Survey respondents ranged from 23-28 students per lab (not all 58 completed every survey); Social desirability bias: Students may have provided favorable feedback due to perception that ChatGPT feedback was provided by instructors; Lack of control group: No comparison with human-generated feedback or no feedback condition; Subjective threshold: Authors acknowledge using an arbitrary 70% threshold for 'favorable' responses without baseline comparison; Selection bias: Only 58 students from single CS1 course at one institution; mostly computer science majors; Attrition: Varying response rates across four surveys (23-28 respondents per survey); Measurement bias: Subjective 70% threshold for determining alignment with formative feedback guidelines (no validated baseline for comparison); Timing bias: Delayed grade release policy limited timely feedback provision; Single-site design reduces generalizability; Selection bias: Only students who consented to participate were included (self-selection); Small sample size (58 students) may not represent broader CS1 student populations; Subjective threshold (70%) for determining 'favorable' alignment with formative feedback guidelines lacks empirical justification; Potential temporal bias from delayed grade release affecting feedback utility perception; No control group for comparison of feedback effectiveness; Survey-based measurement subject to social desirability bias; Limited to single institution (U.S. public university) in summer course
Limitations
- The authors state: "Our study is limited by a small sample of four labs and 58 students, potentially impacting the findings' generalizability
- Future research could expand the participant pool for broader validation." Additionally, "our study primarily examines student perspectives on ChatGPT-generated feedback, but future studies could assess LLM-generated feedback from various angles, including its effects on learning outcomes." The study was "also constrained by the delayed release of grades due to course policy, limiting timely access to feedback," and "to evaluate the extent to which students view ChatGPT-generated feedback as formative, we solely look at the absolute percentage of students who provided favorable responses, using a subjective threshold of 70%
- The lack of a baseline for comparison might limit the conclusiveness of our assessment."
Open questions raised
- Need for studies examining effects of LLM-generated feedback on actual learning outcomes, not just student perceptions
- Development of intuitive software interfaces for educators to generate, evaluate, and optimize LLM-generated feedback end-to-end
- Exploration of incorporating correct code answers into prompts for more efficient error identification
- Fine-tuning LLMs on human-generated feedback and common student errors to provide more specific and corrective feedback
- Improving model understanding of students' prior actions and specific needs
- Examination of settings where feedback is more immediately available (addressing delayed grade release constraint)
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations