RelianceScope: An Analytical Framework for Examining Students' Reliance on Generative AI Chatbots in Problem Solving
Hyoungwook Jin, Minju Yoo, Jieun Han, Zixin Chen, So-Yeon Ahn, Xu Wang · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods design science study with a proof-of-concept application.
Primary method
Design science with iterative framework development. The framework was designed with three explicit design goals: fine-grained analysis, contextual sensitivity, and non-intrusive data collection. Framework components were operationalized through a custom web-based learning system.
Main result
Results show that "active help-seeking is associated with active response-use" (Somers' D test, D = .092, p = .024), and the most prevalent reliance pattern was Passive_Passive, accounting for "44.0% of the 427 interaction segments." Additionally, "reliance patterns remain similar across knowledge mastery levels," and "large language models can reliably detect reliance during help-seeking and response-use" with particularly high accuracy for passive help-seeking (F1 = .807) and response-use (F1 = .760).
Research paradigm
Pragmatist/mixed-methods (design science with empirical validation)
Author conclusions
"We conclude by discussing the implications of RelianceScope and the design guidelines for AI-supported educational systems." The authors state that "the fine-grained, context-sensitive, and non-intrusive design of RelianceScope will inform future investigations of reliance and guide the design of educational systems that better support self-regulated learners." They also note that "the interaction model as a lens to characterize students' challenges in managing reliance during help-seeking and response-use" reveals that "students often struggled to articulate their knowledge gaps and to adapt AI responses."
Risk of bias
Selection bias: Computer science/engineering majors only (limited generalizability); Hawthorne effect: Students aware interactions were being logged; Confounding: Time constraints and grading policies may have influenced reliance behaviors; Measurement bias: Inter-rater reliability for response-use only 71.6% (lower than other measures); Single assessment: One-time post-test limits ability to measure learning retention; Selection bias: Study participants were undergraduate computer science/engineering majors in a single institution, limiting generalizability; Awareness bias: Students knew their interactions were being logged, potentially affecting natural behavior; Demand characteristics: Activity was graded, which may have influenced reliance behaviors; Limited outcome measures: One-time post-test assessment limits ability to measure learning gains robustly; Self-report bias: Reliance on self-reported self-regulation measures and survey responses; Selection bias: Participants were computer science majors in an introductory course, limiting generalizability; Hawthorne effect: Students aware their interactions were logged may have altered behavior; Grading policy effects: Participation was part of course assignment affecting grades, potentially biasing reliance patterns; Single knowledge domain: Limited to web programming (Vue.js), unclear if patterns generalize to other subjects; Inter-rater reliability concerns: Initial agreement rates relatively modest (71.6% for response-use) despite prioritizing constructive patterns; Measurement validity: One-time post-test assessment with small sample (n=79 after exclusions) limits robust statistical inference; Self-reported measures: Questionnaire data subject to response bias and social desirability bias
Limitations
- "First, students' reliance behaviors in our study may not fully generalize to interactions with commercial chatbots (e.g., ChatGPT), as our system employed customized prompts and students were aware that their interactions were logged
- Although we designed appropriate incentives and clearly communicated learning objectives, time constraints and grading policies may have influenced students' reliance behaviors
- Second, we do not establish strong statistical evidence linking reliance patterns to learning gains
- Noise in self-reported measures and the use of a one-time knowledge assessment limited our ability to fit robust regression models."
Open questions raised
- Authors identify need for: (1) applying RelianceScope across varied and controlled learning settings to replicate and extend findings; (2) more diverse learner populations; (3) larger datasets with repeated measures to establish robust relationships between reliance patterns and learning gains; (4) investigation of how reliance patterns generalize to commercial chatbots like ChatGPT; (5) understanding when and how reliance on AI supports learning across different contexts.
- Lack of systematic tools for jointly capturing engagement across help-seeking and response-use in chatbot-assisted learning
- Mixed empirical findings regarding impact of AI reliance on learning, suggesting need for context-sensitive analytical tools
- Need to establish strong statistical links between reliance patterns and learning gains using larger datasets, repeated measures, and diverse learner populations
- Limited understanding of how reliance patterns emerge and affect learning across varied and controlled learning settings
- Need for research examining reliance with commercial chatbots beyond customized systems
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations