Scaffolding Metacognition in Programming Education: Understanding Student-AI Interactions and Design Implications
Boxuan Ma, Huiyong Li, Gen Li, Li Chen, Cheng Tang, Yinjie Xie et al. · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2511.04144
Methodology & findings
Study design
Multi-method study combining: (1) thematic analysis of 2,782 student-AI dialogue logs (26% stratified random sample from 10,632 total interactions) collected over three years from four course offerings; (2) post-course student surveys capturing perceptions of AI support (n=248 students); (3) educator interviews on pedagogical implications.
Sample
N = 248, 4 groups
Primary method
Inter-rater reliability: Cohen's Kappa and percentage agreement. Sequential analysis: Sankey diagrams for visualizing prompt flows. Markov chain analysis for phase transition probabilities. Stratified random sampling (26% of 10,632 dialogue logs = 2,782 samples). Thematic analysis using inductive coding approach with preliminary codebook development on 800 interactions, independent coding on 150 samples with instructor consensus refinement, reliability testing on 334 samples, then independent coding of remaining 2,448 interactions.
Main result
The study found that "Monitoring (M) accounted for the largest share of prompts, far exceeding Planning (P) and Evaluation (E)," indicating that "students most frequently sought AI assistance during active debugging, troubleshooting, and correctness verification, rather than during initial strategy formulation or reflective evaluation." Additionally, "a substantial share of sessions began directly in Monitoring rather than Planning, indicating that students often approached AI reactively after encountering execution problems rather than proactively during problem framing."
Reports effect sizes.
Research paradigm
pragmatic mixed-methods (interpretive + empirical)
Author conclusions
The authors conclude that "AI should not be conceived merely as a source of answers, but as a partner that helps learners navigate metacognitive cycles, develop prompting strategies, and engage with adaptive scaffolding." They synthesize their findings stating: "Synthesizing evidence across these sources, we derived high-level design considerations, highlighting key trade-offs that characterize the emerging design space of educational AI tools," and note that "these findings and design implications will help guide" future development of educational AI tools that "strengthen students' learning processes in programming education."
Risk of bias
Selection bias: Study limited to one institution and single programming language (Python); international student representation only 14.5%; Model confounding: Different GPT versions across cohorts (GPT-3.5-turbo in 2023 vs. GPT-4/4o in later years) may confound results; Attrition/compliance: Participation was voluntary; no information on non-participation rates; Measurement bias: Dialogue logs capture only AI-mediated interactions; offline metacognitive activities not captured; Coder bias: Although inter-rater reliability reported, two coders conducted analysis; potential for shared biases; Social desirability: Survey responses may reflect socially desirable responses rather than actual beliefs; Single institutional setting (Japan-based university) limits generalizability; Self-selection bias: participation was voluntary and optional; Model confound: Different LLM versions (GPT-3.5-turbo in 2023, GPT-4 and GPT-4o in later years) may have influenced interaction patterns and student perceptions; Cohort composition variation across years; Single-modality evidence: Only chat logs captured; offline metacognitive processes not recorded; Researcher coding subjectivity despite inter-rater reliability measures; Single-institution sampling (may not generalize to other universities or countries); Selection bias: participation was voluntary and not mandatory; LLM model variation across years (GPT-3.5-turbo in 2023 vs. GPT-4/GPT-4o in later years) may confound findings; Observable interaction bias: students may adjust behavior knowing interactions are logged; Multi-modal activity not captured: peer discussions, IDE work, and consultation of learning materials outside chatbot logs
Limitations
- "Our findings derive from a single Python programming course at one institution, which constrains the breadth of applicability
- Validation across courses in other languages (such as C or Java), as well as across different institutional contexts and cultural environments, will be important for assessing the robustness of these results." Additionally, "this study examined only student–LLM interactions, however students may expressing their critical thinking in ways that are not captured in chatbot logs
- In actual classroom contexts, learning unfolds across multiple modalities, such as seeking help from peers or instructors, working within IDEs, consulting learning materials, or referring back to prior exercise solutions."
Open questions raised
- Limited cross-institutional validation; need for validation across courses in other languages (C, Java) and different institutional contexts
- Model versioning effects: need to understand how different LLM versions shape interaction patterns and student perceptions
- Multi-modal learning: unclear how chatbot interactions integrate with peer learning, instructor support, IDE use, and learning materials
- Individual differences: limited investigation of how learner profiles (self-efficacy, prior knowledge) shape metacognitive engagement with AI
- Scaffolding forms: pedagogical potential of fill-in-the-blank code, Socratic questioning, and other intermediate scaffolds remains understudied
- Limited understanding of how metacognitive processes evolve when learners interact with AI-powered assistants
Explore related topics
Related papers
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- Revolutionizing education with AI: Exploring the transformative potential of ChatGPTTufan Adıgüzel · 2023 · 858 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations