12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Multimodal Methods for Analyzing Learning and Training Environments: A Systematic Literature Review

Clayton Cohn, Eduardo Davalos, Caleb Vatral, Joyce Horn Fonteles, Hanchen David Wang · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
2/4
Quality (LMQS)
I
Evidence
1
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2408.14491

Methodology & findings

Study design

Systematic literature review following Kitchenham's systematic review methodology.

Main result

The study found that "integrating modalities enables richer insights into learner and trainee behaviors, revealing latent patterns often overlooked by unimodal approaches." The analysis identified five modality groups (Natural Language, Vision, Physiological Signals, Human-Centered Evidence, and Environment Logs) and revealed that physiological signal modalities appear more frequently in post-LLM research (Corpus B: 23/49; 47%) compared to pre-LLM work (Corpus A: 20/73; 27%), with a marked shift away from vision modalities in recent years—from 59/73 (81%) in Corpus A to 27/49 (55%) in Corpus B. The corpus shows that "human-centered data are most often used for model-free, qualitative analysis" and that LLMs have broadened the interpretive scope by enabling direct integration of log data into prompts as contextualized natural language.

Research paradigm

Mixed methods (qualitative thematic analysis with quantitative frequency counts)

Author conclusions

The authors conclude that multimodal learning analytics requires fundamental shifts in methodology to address real-world constraints. They state: "Applied MMLA differs fundamentally from multimodal research in core AI and machine learning. In real-world settings, data must be collected under conditions that introduce noise, partial observability, missing data, and privacy constraints." They emphasize that "a comprehensive synthesis of applied multimodal methods is, therefore, needed to support the design, implementation, and interpretation of multimodal analyses in learning and training contexts." The review demonstrates that "integrating modalities enables richer insights into learner and trainee behaviors, revealing latent patterns often overlooked by unimodal approaches," but that "persistent challenges in multimodal data collection and integration continue to hinder the adoption of these systems in real-time classroom settings." The authors advocate for recognizing the bifurcation between GenAI and non-GenAI approaches, stating: "As a result, the literature increasingly bifurcates into GenAI and non-GenAI approaches, with each operating under different assumptions about model capability, data availability, and the role of automation in analytic workflows. Recognizing and unpacking this divide is essential for contextualizing current methods within the rapidly evolving MMLA landscape."

Risk of bias

Selection bias: Google Scholar search may not capture all relevant literature; potential bias toward English-language publications; Citation graph pruning as filtering heuristic may exclude relevant papers with weak citation connections; Reviewer bias: qualitative screening by two authors may not eliminate subjective judgment in inclusion/exclusion decisions; Publication bias: systematic review does not employ funnel plot analysis or publication bias assessment; Domain bias: corpus heavily skewed toward STEM+C (75% Corpus A, 76% Corpus B) and university learners (49% A, 59% B); Temporal bias: pre-ChatGPT corpus (2017-2022) contains 73 papers while post-ChatGPT (2022-2025) contains 49, reflecting different time spans and relative recency of field; Selection bias: Initial search of 2,120 papers (Corpus A) and 845 (Corpus B) reduced through citation graph pruning, which may exclude relevant papers with weaker citation networks; Language bias: Non-English papers excluded from dataset; Publication bias: Review limited to published peer-reviewed and preprint literature; gray literature not included; Database bias: Google Scholar via SerpAPI used as primary search source; may not capture all relevant databases; Reviewer bias: Manual qualitative screening by at least two authors introduces potential coding inconsistencies despite dual review; Domain bias: Corpus heavily skewed toward STEM+C (75-76%) and university-level instruction (49-59%), underrepresenting humanities and professional development; Recency bias: Corpus B deliberately selected post-ChatGPT papers, potentially overrepresenting LLM-based approaches; Selection bias: Search limited to Google Scholar and English-language papers only; Publication bias: Review restricted to published studies, excluding grey literature; Author bias: Citation graph pruning is pragmatic filtering heuristic that may exclude relevant papers; Screening bias: Potential for reviewer disagreement despite dual-review process; Domain bias: Corpus heavily skewed toward STEM+C (75-76% across corpora) and university-level instruction (49-59%)

Limitations

  • The review explicitly states that "a comprehensive review of empirical methods in applied multimodal environments remains notably absent" prior to this work
  • The scope is limited in several ways: "This review focuses on multimodal learning and training analytics in empirical studies conducted in authentic instructional and training settings
  • We exclude studies centered exclusively on virtual reality environments, as they face scalability constraints for applied educational deployment." The review acknowledges that "prior surveys primarily describe fusion taxonomies, conceptual frameworks, methodological overviews, cross-domain multimodal machine learning, or domain-specific applications like healthcare or sensing
  • They do not synthesize how methodological decisions are shaped by the constraints of real-world data collection." Additionally, the authors note methodological challenges persist: "many learning and training environments lack controlled lighting, fixed camera setups, or specialized hardware (e.g., eye trackers), limiting the feasibility of fine-grained gaze or pose analysis."

Open questions raised

  • Lack of synthesis on how methodological decisions are shaped by real-world data collection constraints (noise, partial observability, missing data, privacy)
  • Limited examination of cross-modal interactions and temporal alignment challenges in heterogeneous multimodal integration
  • Insufficient analysis of empirical practices in applied settings (classrooms, clinical simulations, workplace training) versus laboratory conditions
  • Underexplored methodological consequences of integrating modalities under real-world constraints
  • Missing grounding on how recent advances in LLMs and GenAI reshape multimodal data collection, analysis, and system design
  • Limited evidence-derived taxonomy of multimodal methods in applied learning and training contexts
Data: No datasets mentioned as available. The review is a synthesis of 122 published papers.Code: None mentionedExtracted from: pdfAgreement 64%

Explore related topics

Related papers