Three Years with Classroom AI in Introductory Programming: Shifts in Student Awareness, Interaction, and Performance
Boxuan Ma, Huiyong Li, Gen Li, Li Chen, Cheng Tang, Atsushi Shimada et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Longitudinal mixed-methods classroom study spanning three years (2023-2025) with three successive cohorts in an introductory Python course.
Sample
N = 248, 11 groups
Primary method
Kruskal-Wallis H test for comparing interaction frequency and assignment performance across cohorts (non-parametric alternative to ANOVA); Cohen's Kappa for inter-rater reliability assessment of thematic coding; Thematic analysis of dialogue logs using stratified random sampling and consensus coding; Transition Network Analysis (TNA) for sequential patterns in student prompts; Descriptive statistics (means, standard deviations) for cohort-level comparisons; Percentage agreement for coding reliability
Main result
The study found that "students' relationships with GenAI change systematically over time: familiarity and uptake become increasingly normative, and help-seeking practices evolve alongside growing AI literacy and shifting expectations of what the assistant should provide." Additionally, "assignment scores were consistently high and tightly clustered across cohorts" with "no statistically significant differences across cohorts (H = 3.23, p = 0.357)," suggesting that "AI access alone did not substantially change observable performance on course assignments or overall grades." The research also demonstrates that "students increasingly treat the assistant as a partner for verification and iterative refinement rather than a one-shot code generator," with interaction patterns shifting from "implementation-driven style (2023), to a verification-driven style (2024), and then to a more balanced debug–verify–implement pattern (2025)."
Reports effect sizes.
Research paradigm
Mixed-methods empirical (quantitative + qualitative)
Author conclusions
The authors conclude: "Across cohorts, students entered with steadily higher GenAI familiarity and more routine use. Dialogue logs show a shift from implementation-first, short exchanges to verification-centered, multi-step workflows. Despite these changes in interaction strategies, cohort-level assignment and final-grade outcomes were broadly comparable. These findings provide a longitudinal baseline for future work on designing and evaluating instructional supports that foster effective AI literacy and responsible help-seeking in programming learning. Beyond documenting adoption, we show that as GenAI becomes normalized, students increasingly treat the assistant as a partner for verification and iterative refinement rather than a one-shot code generator. This suggests that the central instructional challenge is not simply whether students use GenAI, but how courses shape productive use and align help-seeking with course expectations."
Risk of bias
Selection bias: Single institution, specific programming course context; Confounding variables: Cannot isolate GenAI improvements from model quality improvements or peer norms across cohorts; Attrition/missing data: Not explicitly discussed; Measurement bias: Dialogue logs only capture in-platform use, missing alternative help-seeking channels (peers, TAs, external resources); Social desirability bias: Self-reported familiarity and usage in questionnaires; Selection bias: Voluntary participation; self-selection effects in GenAI adoption across cohorts; Confounding: Cohort-level design cannot isolate GenAI effects from broader technological changes, model improvements, or peer norm shifts; Measurement limitation: Dialogue logs capture only in-platform interactions, missing alternative help-seeking channels (peers, TAs, external resources); Attrition/missing data: No explicit reporting of dropout rates or missing questionnaire responses; Grading scheme effects: Multiple-submission policy may have buffered marginal effects of GenAI access on grades; Selection bias: Voluntary participation may skew toward students more motivated to use AI; Confounding by cohort effects: Cannot separate GenAI model improvements from student behavior changes; Measurement bias: Dialogue logs capture only in-platform use, missing alternative help-seeking channels; Attrition: No mention of dropout rates or missing data across years; Grading scheme confound: Multiple submission attempts before grading may obscure learning differences; Cohort composition shifts: Different group sizes (62→126→60) and enrollment changes across years
Limitations
- The authors state: "This study has several limitations
- First, it focuses on a single institution and course context, so patterns may differ across populations, programming languages, or grading schemes
- Second, our cohort-level observational design cannot establish causality (e.g., model improvements or peer norms)
- Third, dialogue logs capture only in-platform GenAI use and may miss other support channels (e.g., peers or TAs)."
Open questions raised
- Limited longitudinal evidence on how students' awareness of AI, student-AI interaction patterns, and course outcomes evolve as AI becomes routine in classrooms
- Lack of classroom evidence jointly tracking prior familiarity and expectations, student-AI interaction pattern shifts, and coinciding changes in course outcomes
- Need for designing instructional supports that foster effective AI literacy and responsible help-seeking
- Understanding how courses redefine productive learning practices while maintaining student agency in the AI era
- Longitudinal evidence on how students' awareness, interaction patterns, and outcomes evolve as AI becomes routine in classrooms (addressed by this study)
- Need for causal evidence on GenAI effects independent of model improvements and peer norms
Explore related topics
Related papers
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations