12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

RelianceScope: An Analytical Framework for Examining Students' Reliance on Generative AI Chatbots in Problem Solving

Hyoungwook Jin, Minju Yoo, Jieun Han, Zixin Chen, So-Yeon Ahn, Xu Wang · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods design science study with a proof-of-concept application.

Primary method

Design science with iterative framework development. The framework was designed with three explicit design goals: fine-grained analysis, contextual sensitivity, and non-intrusive data collection. Framework components were operationalized through a custom web-based learning system.

Main result

Results show that "active help-seeking is associated with active response-use" (Somers' D test, D = .092, p = .024), and the most prevalent reliance pattern was Passive_Passive, accounting for "44.0% of the 427 interaction segments." Additionally, "reliance patterns remain similar across knowledge mastery levels," and "large language models can reliably detect reliance during help-seeking and response-use" with particularly high accuracy for passive help-seeking (F1 = .807) and response-use (F1 = .760).

Research paradigm

Pragmatist/mixed-methods (design science with empirical validation)

Author conclusions

"We conclude by discussing the implications of RelianceScope and the design guidelines for AI-supported educational systems." The authors state that "the fine-grained, context-sensitive, and non-intrusive design of RelianceScope will inform future investigations of reliance and guide the design of educational systems that better support self-regulated learners." They also note that "the interaction model as a lens to characterize students' challenges in managing reliance during help-seeking and response-use" reveals that "students often struggled to articulate their knowledge gaps and to adapt AI responses."

Risk of bias

Selection bias: Computer science/engineering majors only (limited generalizability); Hawthorne effect: Students aware interactions were being logged; Confounding: Time constraints and grading policies may have influenced reliance behaviors; Measurement bias: Inter-rater reliability for response-use only 71.6% (lower than other measures); Single assessment: One-time post-test limits ability to measure learning retention; Awareness bias: Students knew their interactions were being logged, potentially affecting natural behavior; Demand characteristics: Activity was graded, which may have influenced reliance behaviors; Self-report bias: Reliance on self-reported self-regulation measures and survey responses; Grading policy effects: Participation was part of course assignment affecting grades, potentially biasing reliance patterns; Single knowledge domain: Limited to web programming (Vue.js), unclear if patterns generalize to other subjects; Inter-rater reliability concerns: Initial agreement rates relatively modest (71.6% for response-use) despite prioritizing constructive patterns

Limitations

  • "First, students' reliance behaviors in our study may not fully generalize to interactions with commercial chatbots (e.g., ChatGPT), as our system employed customized prompts and students were aware that their interactions were logged
  • Although we designed appropriate incentives and clearly communicated learning objectives, time constraints and grading policies may have influenced students' reliance behaviors
  • Second, we do not establish strong statistical evidence linking reliance patterns to learning gains
  • Noise in self-reported measures and the use of a one-time knowledge assessment limited our ability to fit robust regression models."

Open questions raised

  • Authors identify need for:
  • applying RelianceScope across varied and controlled learning settings to replicate and extend findings
  • more diverse learner populations
  • larger datasets with repeated measures to establish robust relationships between reliance patterns and learning gains
  • investigation of how reliance patterns generalize to commercial chatbots like ChatGPT
  • understanding when and how reliance on AI supports learning across different contexts.
Data: RelianceScope annotated dataset with 1,362 chat logs, 2,708 code edit logs, pre/post-assessments, and self-regulation measures from 79 students: https://osf.io/27ec5/overview?view_only=a8234a17f908464297d35d5ca1ef476cCode: Not mentionedExtracted from: pdf

Explore related topics

Related papers