12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Students’ Perceptions and Preferences of Generative Artificial Intelligence Feedback for Programming

Zhengdong Zhang, Zihan Dong, Yang Shi, Thomas Price, Noboru Matsuda, Dongkuan Xu · Proceedings of the AAAI Conference on Artificial Intelligence · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
3/4
Quality (LMQS)
E
Evidence
37
Citations
13.14
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30372

Methodology & findings

Study design

Survey-based empirical study with thematic qualitative analysis.

Sample

N = 58, 5 groups

Primary method

Descriptive statistics (percentages, frequencies) calculated for Likert-scale survey responses. Thematic analysis conducted independently by three researchers with collaborative reconciliation of codes. No inferential statistical tests (e.g., t-tests, ANOVA, chi-square) reported. Percentages calculated for preference data and guideline alignment (e.g., 72.5%, 98%, 95%, 71%).

Main result

The study found that "the ChatGPT-generated feedback largely aligns with formative feedback guidelines, with over 70% of students giving favorable evaluations of this alignment." Additionally, "72.5% (74/102) of the responses preferred the feedback generated with the prompt that contains students' code (feedback 1) over the feedback without students' code (feedback 2)," with students citing specificity (33/74), clarity (15/74), and corrective suggestions (9/74) as the top three reasons for their preference.

Reports effect sizes.

Research paradigm

Empirical-interpretivist (mixed methods: survey + thematic analysis)

Author conclusions

"This study demonstrated that ChatGPT could generate Java programming assignment feedback that students perceived as formative. It also offered insights into the specific improvements that would make the ChatGPT-generated feedback useful for students." The authors note that "ChatGPT has considerable potential for generating practical and meaningful assignment feedback" and recommend that "educators could explore personalization options to better meet students' diverse needs in real-world educational contexts," as "student preferences for feedback tone varied. Some students favored a positive tone, while others sought more critical assessments."

Risk of bias

Selection bias: All participants were from a single introductory Java course at one U.S. public university, predominantly computer science majors; Attrition: Survey respondents ranged from 23-28 students per lab (not all 58 completed every survey); Social desirability bias: Students may have provided favorable feedback due to perception that ChatGPT feedback was provided by instructors; Lack of control group: No comparison with human-generated feedback or no feedback condition; Subjective threshold: Authors acknowledge using an arbitrary 70% threshold for 'favorable' responses without baseline comparison; Selection bias: Only 58 students from single CS1 course at one institution; mostly computer science majors; Attrition: Varying response rates across four surveys (23-28 respondents per survey); Measurement bias: Subjective 70% threshold for determining alignment with formative feedback guidelines (no validated baseline for comparison); Timing bias: Delayed grade release policy limited timely feedback provision; Single-site design reduces generalizability; Selection bias: Only students who consented to participate were included (self-selection); Small sample size (58 students) may not represent broader CS1 student populations; Subjective threshold (70%) for determining 'favorable' alignment with formative feedback guidelines lacks empirical justification; Potential temporal bias from delayed grade release affecting feedback utility perception; No control group for comparison of feedback effectiveness; Survey-based measurement subject to social desirability bias; Limited to single institution (U.S. public university) in summer course

Limitations

  • The authors state: "Our study is limited by a small sample of four labs and 58 students, potentially impacting the findings' generalizability
  • Future research could expand the participant pool for broader validation." Additionally, "our study primarily examines student perspectives on ChatGPT-generated feedback, but future studies could assess LLM-generated feedback from various angles, including its effects on learning outcomes." The study was "also constrained by the delayed release of grades due to course policy, limiting timely access to feedback," and "to evaluate the extent to which students view ChatGPT-generated feedback as formative, we solely look at the absolute percentage of students who provided favorable responses, using a subjective threshold of 70%
  • The lack of a baseline for comparison might limit the conclusiveness of our assessment."

Open questions raised

  • Need for studies examining effects of LLM-generated feedback on actual learning outcomes, not just student perceptions
  • Development of intuitive software interfaces for educators to generate, evaluate, and optimize LLM-generated feedback end-to-end
  • Exploration of incorporating correct code answers into prompts for more efficient error identification
  • Fine-tuning LLMs on human-generated feedback and common student errors to provide more specific and corrective feedback
  • Improving model understanding of students' prior actions and specific needs
  • Examination of settings where feedback is more immediately available (addressing delayed grade release constraint)
Data: Not mentioned as publicly availableCode: Not mentionedExtracted from: pdfAgreement 62%

Explore related topics

Related papers