Computer Science Education in ChatGPT Era: Experiences from an Experiment in a Programming Course for Novice Programmers
Tomaž Kosar, Dragana Ostojić, Yu David Liu, Marjan Mernik · Mathematics · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/math12050629
Methodology & findings
Study design
Controlled between-subjects experiment with two groups (treatment group using ChatGPT, control group without ChatGPT) of 182 first-year undergraduate students.
Sample
N = 182, 6 groups
Primary method
Shapiro-Wilk test of normal distribution for all data. If data were not normally distributed, Mann-Whitney U test (non-parametric) for two independent samples was performed. For normally distributed data, Independent Sample t-test was used. Statistical significance threshold: α = 0.05.
Main result
The study found that "the students' performance is not influenced by ChatGPT usage (no statistical significance between groups with a p-value of 0.730), nor are the grading results of practical assignments (p-value 0.760) and midterm exams (p-value 0.856)." The authors conclude that "using LLM is not a decisive factor if the right actions are taken before the execution of the course" and that their controlled experiment suggests "it is safe for novice programmers to use ChatGPT if specific measures and adjustments are adopted in the education process."
Reports effect sizes.
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude: "The main findings suggest the following: Comparing the participants' success in practical assignments between groups using ChatGPT and others not using it, we found that the results were not statistically different... Comparing the participants' success in midterm exams between groups using ChatGPT and others not using it, we found that the results were also not statistically different... Comparing the participants' overall success in a course on Programming II between groups using ChatGPT and others not using it, we found that the results were also not statistically different. This means that our specific execution of the course (with all the introduced adjustments) allows using ChatGPT as an additional learning aid."
Risk of bias
Selection bias: Participants were from a single institution (University of Maribor, Slovenia), attracting 'the best students from the country,' limiting generalizability; Attrition: 16 participants eliminated (198 to 182) for not completing assignments or midterm exams; Procedural bias: 8 students removed from Group II who reported ChatGPT use, creating potential selection bias; Measurement bias: Paper-based midterms may not capture performance under real-world coding conditions with IDE support; Confounding: Assignment defense procedures and carefully constructed assignments may differentially benefit both groups regardless of ChatGPT access; Hawthorne effect: Students aware of being in an experiment may alter behavior; Gender imbalance: 85.9% male, 12.4% female, 1.7% preferred not to say; Selection bias: Only one institution (University of Maribor) with best students in the country; Attrition: Started with 198 participants, excluded 16 for non-completion of assignments or midterm exams; Measurement validity: Single course, single programming language (C++); Confounding: No explicit control for prior ChatGPT experience beyond baseline Likert scale measurement; Hawthorne effect: Students aware they were in an experiment; Contamination: Eight students from control group (Group II) reported using ChatGPT and were removed post-hoc; Selection bias: Single institution (University of Maribor), may not generalize to other contexts; Attrition: Excluded 16 participants (198 to 182) for not completing assignments or exams; Confounding: Removed 8 students from Group II who self-reported ChatGPT use, potentially biasing group composition; Demand characteristics: Students aware of experiment, may alter behavior; Instructor expectancy effects: Lecturers/TAs aware of group assignments despite efforts to keep blinded; Single programming language: Only C++ used, limits generalizability; Hawthorne effect: Intensive defense procedures and monitoring may change natural behavior
Limitations
- The authors state that "Our study results and ChatGPT-oriented adjustments must be taken with caution in the future
- Improvements in large language models will likely affect adjustments (specifically for practical assignments)." Additionally, they note that "This study needs additional replications
- Different problems (applications) need to be applied to lab work with a different programming language, to name a few possibilities for strengthening the validity of our conclusions." They also acknowledge that "broadening the perspective in replicated studies to involve more institutions and conducting a multi-institutional and multinational study has the potential to yield a deeper comprehension of the integration of large language models in education."
Open questions raised
- Need for replication studies with different programming languages and applications
- Comparison of results from midterm exams using IDE support versus paper-based format
- Application of experiment design to introductory programming (CS1) courses
- Empirical study to understand essential skills (critical thinking, problem-solving, group work skills)
- Qualitative assessments to complement performance metrics and uncover cognitive engagement
- Investigation of identified risks: unreliability of generated data, students' reliance on technology, potential impact on cognitive abilities and interpersonal communication
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations