Using Large Language Models to Detect Socially Shared Regulation of Collaborative Learning
Jiayi Zhang, Conrad Borchers, Clayton Cohn, Namrata Srivastava, Caitlin Snyder, Siyuan Guo et al. · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1145/3785022.3785083
Methodology & findings
Study design
Computational predictive modeling study using multimodal data integration.
Main result
The study found that "text_only embeddings are strongest on average" with a mean AUC of 0.650 [0.615, 0.685], and "text embeddings excel for enactment-oriented phases (ENACTING: 0.6745; ENACTING & PLANNING: 0.7285; ENACTING & MONITORING: 0.6796)," while "short-range conversational context is particularly valuable for PLANNING (utterance: 0.5036 →text_with_context: 0.6711; multimodal: 0.4803 →0.6269)." Overall, the detectors achieved "practically useful accuracy for several SRL behaviors, with complementary strengths: multimodal traces (text+logs) are most discriminative when behaviors align with concrete environment use."
Research paradigm
Positivist/empiricist with pragmatist elements
Author conclusions
The authors conclude that "this work demonstrates the feasibility of using embedding-based deep learning to detect SSRL behaviors and group dynamics in collaborative STEM+C learning environments by integrating discourse and log data aligned with task context. By systematically comparing text-only, contextualized, and multimodal representations, we show that discourse provides a strong foundation for modeling, while context and log features offer complementary value for specific constructs. Beyond advancing scalable SSRL detection, our approach highlights the potential for multimodal learning analytics to provide educators with actionable insights into group regulation processes, ultimately supporting more effective feedback and intervention in collaborative classrooms."
Risk of bias
Selection bias: Participants were from structured university-affiliated program, not representative of broader student populations; Small sample size: 36 high school sophomores limits generalizability; Task-specific models: Trained on single Truck Task in C2STEM environment; Demographic data not collected: Race and gender information unavailable, preventing analysis of potential disparate impacts; Teacher feedback: Only 4 teachers, selection method not specified; Manual coding by authors: Inter-rater reliability achieved but potential for systematic bias in how research team interprets segments; Selection bias: participants were from a specific NSF-supported program affiliated with Vanderbilt University, not representative of broader student populations; Small sample size: only 36 high school sophomores across two studies; Single task environment: models trained only on C2STEM Truck Task, limiting generalizability; Label imbalance: code distributions ranged from 7% to 22%, addressed through class-balanced weights; Manual verification bias: two researchers manually corrected transcriptions and verified summaries, potential for inter-rater drift; Selection bias: convenience sample of 36 students from affiliated Vanderbilt program; Limited demographic diversity: demographic data on race and gender could not be collected; Single environment bias: models trained on single task (Truck Task) in single learning environment (C2STEM); Single task generalization risk: results may not transfer to other computational modeling tasks; Coder bias: manual coding by research team members, though inter-rater reliability was established
Limitations
- The authors state that "as the results indicate, different constructs benefit from different sources of information
- some are captured well through discourse alone, while others rely more on contextual cues or multimodal data
- Although the best-performing models achieved reasonable accuracy, there remains substantial room for improvement." They also acknowledge that "the models were trained on data from students collaborating on a single task within a single learning environment" and note that "future work should evaluate the transferability of this approach across subject domains, task types, and broader student populations to ensure its robustness in diverse classroom settings."
Open questions raised
- Generalizability across different tasks, domains, and student populations
- Need for more advanced architectures (hierarchical models accounting for discourse flow)
- Incorporation of richer multimodal features (gaze, gestures)
- Designing analytics that are interpretable and actionable for teachers
- Addressing teacher concerns about multiple labels within single segments
- Exploring visual presentation formats (dashboards, multimodal timelines) for practitioner use
Explore related topics
Related papers
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Revolutionizing education with AI: Exploring the transformative potential of ChatGPTTufan Adıgüzel · 2023 · 858 citations