Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review
Clayton Cohn, Eduardo Davalos, Caleb Vatral, Joyce Horn Fonteles, Hanchen David Wang · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2408.14491
Methodology & findings
Study design
Systematic literature review following Kitchenham's systematic review methodology.
Main result
The study found that "integrating modalities enables richer insights into learner and trainee behaviors, revealing latent patterns often overlooked by unimodal approaches." The analysis identified five modality groups (Natural Language, Vision, Physiological Signals, Human-Centered Evidence, and Environment Logs) and revealed that physiological signal modalities appear more frequently in post-LLM research (Corpus B: 23/49; 47%) compared to pre-LLM work (Corpus A: 20/73; 27%), with a marked shift away from vision modalities in recent years—from 59/73 (81%) in Corpus A to 27/49 (55%) in Corpus B. The corpus shows that "human-centered data are most often used for model-free, qualitative analysis" and that LLMs have broadened the interpretive scope by enabling direct integration of log data into prompts as contextualized natural language.
Research paradigm
Mixed methods (qualitative thematic analysis with quantitative frequency counts)
Author conclusions
The authors conclude that multimodal learning analytics requires fundamental shifts in methodology to address real-world constraints. They state: "Applied MMLA differs fundamentally from multimodal research in core AI and machine learning. In real-world settings, data must be collected under conditions that introduce noise, partial observability, missing data, and privacy constraints." They emphasize that "a comprehensive synthesis of applied multimodal methods is, therefore, needed to support the design, implementation, and interpretation of multimodal analyses in learning and training contexts." The review demonstrates that "integrating modalities enables richer insights into learner and trainee behaviors, revealing latent patterns often overlooked by unimodal approaches," but that "persistent challenges in multimodal data collection and integration continue to hinder the adoption of these systems in real-time classroom settings." The authors advocate for recognizing the bifurcation between GenAI and non-GenAI approaches, stating: "As a result, the literature increasingly bifurcates into GenAI and non-GenAI approaches, with each operating under different assumptions about model capability, data availability, and the role of automation in analytic workflows. Recognizing and unpacking this divide is essential for contextualizing current methods within the rapidly evolving MMLA landscape."
Risk of bias
Selection bias: Google Scholar search may not capture all relevant literature; potential bias toward English-language publications; Citation graph pruning as filtering heuristic may exclude relevant papers with weak citation connections; Reviewer bias: qualitative screening by two authors may not eliminate subjective judgment in inclusion/exclusion decisions; Publication bias: systematic review does not employ funnel plot analysis or publication bias assessment; Domain bias: corpus heavily skewed toward STEM+C (75% Corpus A, 76% Corpus B) and university learners (49% A, 59% B); Temporal bias: pre-ChatGPT corpus (2017-2022) contains 73 papers while post-ChatGPT (2022-2025) contains 49, reflecting different time spans and relative recency of field; Selection bias: Initial search of 2,120 papers (Corpus A) and 845 (Corpus B) reduced through citation graph pruning, which may exclude relevant papers with weaker citation networks; Language bias: Non-English papers excluded from dataset; Publication bias: Review limited to published peer-reviewed and preprint literature; gray literature not included; Database bias: Google Scholar via SerpAPI used as primary search source; may not capture all relevant databases; Reviewer bias: Manual qualitative screening by at least two authors introduces potential coding inconsistencies despite dual review; Domain bias: Corpus heavily skewed toward STEM+C (75-76%) and university-level instruction (49-59%), underrepresenting humanities and professional development; Recency bias: Corpus B deliberately selected post-ChatGPT papers, potentially overrepresenting LLM-based approaches; Selection bias: Search limited to Google Scholar and English-language papers only; Publication bias: Review restricted to published studies, excluding grey literature; Author bias: Citation graph pruning is pragmatic filtering heuristic that may exclude relevant papers; Screening bias: Potential for reviewer disagreement despite dual-review process; Domain bias: Corpus heavily skewed toward STEM+C (75-76% across corpora) and university-level instruction (49-59%)
Limitations
- The review explicitly states that "a comprehensive review of empirical methods in applied multimodal environments remains notably absent" prior to this work
- The scope is limited in several ways: "This review focuses on multimodal learning and training analytics in empirical studies conducted in authentic instructional and training settings
- We exclude studies centered exclusively on virtual reality environments, as they face scalability constraints for applied educational deployment." The review acknowledges that "prior surveys primarily describe fusion taxonomies, conceptual frameworks, methodological overviews, cross-domain multimodal machine learning, or domain-specific applications like healthcare or sensing
- They do not synthesize how methodological decisions are shaped by the constraints of real-world data collection." Additionally, the authors note methodological challenges persist: "many learning and training environments lack controlled lighting, fixed camera setups, or specialized hardware (e.g., eye trackers), limiting the feasibility of fine-grained gaze or pose analysis."
Open questions raised
- Lack of synthesis on how methodological decisions are shaped by real-world data collection constraints (noise, partial observability, missing data, privacy)
- Limited examination of cross-modal interactions and temporal alignment challenges in heterogeneous multimodal integration
- Insufficient analysis of empirical practices in applied settings (classrooms, clinical simulations, workplace training) versus laboratory conditions
- Underexplored methodological consequences of integrating modalities under real-world constraints
- Missing grounding on how recent advances in LLMs and GenAI reshape multimodal data collection, analysis, and system design
- Limited evidence-derived taxonomy of multimodal methods in applied learning and training contexts
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations