12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Computer Science Education in ChatGPT Era: Experiences from an Experiment in a Programming Course for Novice Programmers

Tomaž Kosar, Dragana Ostojić, Yu David Liu, Marjan Mernik · Mathematics · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
3/4
Quality (LMQS)
E
Evidence
81
Citations
24.38
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3390/math12050629

Methodology & findings

Study design

Controlled between-subjects experiment with two groups (treatment group using ChatGPT, control group without ChatGPT) of 182 first-year undergraduate students.

Sample

N = 182, 6 groups

Primary method

Shapiro-Wilk test of normal distribution for all data. If data were not normally distributed, Mann-Whitney U test (non-parametric) for two independent samples was performed. For normally distributed data, Independent Sample t-test was used. Statistical significance threshold: α = 0.05.

Main result

The study found that "the students' performance is not influenced by ChatGPT usage (no statistical significance between groups with a p-value of 0.730), nor are the grading results of practical assignments (p-value 0.760) and midterm exams (p-value 0.856)." The authors conclude that "using LLM is not a decisive factor if the right actions are taken before the execution of the course" and that their controlled experiment suggests "it is safe for novice programmers to use ChatGPT if specific measures and adjustments are adopted in the education process."

Reports effect sizes.

Research paradigm

Positivist/Empiricist

Author conclusions

The authors conclude: "The main findings suggest the following: Comparing the participants' success in practical assignments between groups using ChatGPT and others not using it, we found that the results were not statistically different... Comparing the participants' success in midterm exams between groups using ChatGPT and others not using it, we found that the results were also not statistically different... Comparing the participants' overall success in a course on Programming II between groups using ChatGPT and others not using it, we found that the results were also not statistically different. This means that our specific execution of the course (with all the introduced adjustments) allows using ChatGPT as an additional learning aid."

Risk of bias

Selection bias: Participants were from a single institution (University of Maribor, Slovenia), attracting 'the best students from the country,' limiting generalizability; Attrition: 16 participants eliminated (198 to 182) for not completing assignments or midterm exams; Procedural bias: 8 students removed from Group II who reported ChatGPT use, creating potential selection bias; Measurement bias: Paper-based midterms may not capture performance under real-world coding conditions with IDE support; Confounding: Assignment defense procedures and carefully constructed assignments may differentially benefit both groups regardless of ChatGPT access; Hawthorne effect: Students aware of being in an experiment may alter behavior; Gender imbalance: 85.9% male, 12.4% female, 1.7% preferred not to say; Selection bias: Only one institution (University of Maribor) with best students in the country; Attrition: Started with 198 participants, excluded 16 for non-completion of assignments or midterm exams; Measurement validity: Single course, single programming language (C++); Confounding: No explicit control for prior ChatGPT experience beyond baseline Likert scale measurement; Hawthorne effect: Students aware they were in an experiment; Contamination: Eight students from control group (Group II) reported using ChatGPT and were removed post-hoc; Selection bias: Single institution (University of Maribor), may not generalize to other contexts; Attrition: Excluded 16 participants (198 to 182) for not completing assignments or exams; Confounding: Removed 8 students from Group II who self-reported ChatGPT use, potentially biasing group composition; Demand characteristics: Students aware of experiment, may alter behavior; Instructor expectancy effects: Lecturers/TAs aware of group assignments despite efforts to keep blinded; Single programming language: Only C++ used, limits generalizability; Hawthorne effect: Intensive defense procedures and monitoring may change natural behavior

Limitations

  • The authors state that "Our study results and ChatGPT-oriented adjustments must be taken with caution in the future
  • Improvements in large language models will likely affect adjustments (specifically for practical assignments)." Additionally, they note that "This study needs additional replications
  • Different problems (applications) need to be applied to lab work with a different programming language, to name a few possibilities for strengthening the validity of our conclusions." They also acknowledge that "broadening the perspective in replicated studies to involve more institutions and conducting a multi-institutional and multinational study has the potential to yield a deeper comprehension of the integration of large language models in education."

Open questions raised

  • Need for replication studies with different programming languages and applications
  • Comparison of results from midterm exams using IDE support versus paper-based format
  • Application of experiment design to introductory programming (CS1) courses
  • Empirical study to understand essential skills (critical thinking, problem-solving, group work skills)
  • Qualitative assessments to complement performance metrics and uncover cognitive engagement
  • Investigation of identified risks: unreliability of generated data, students' reliance on technology, potential impact on cognitive abilities and interpersonal communication
Data: https://github.com/tomazkosar/DifferentialStudyChatGPT (accessed on 12 January 2024) - contains background questionnaires, feedback questionnaires, practical assignment descriptions, and experimental data; https://github.com/tomazkosar/DifferentialStudyChatGPT; https://github.com/tomazkosar/DifferentialStudyChatGPT (accessed 12 January 2024)Code: https://github.com/tomazkosar/DifferentialStudyChatGPTExtracted from: pdfAgreement 49%

Explore related topics

Related papers