12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

GAIVA: A Generative AI Framework for Human-AI Collaboration in Educational Video Analysis

Yiqiu Zhou, Jina Kang, Ha Nguyen · Open MIND · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.5281/zenodo.20748282

Methodology & findings

Study design

Case study using a modular computational pipeline (GAIVA) integrating multimodal large language models (MLLMs) and large language models (LLMs) for qualitative video analysis.

Main result

The study found that "the role-assignment achieved the most balanced performance in accuracy (M = 3.24, SD = 1.72), completeness (M = 3.86, SD = 0.35), and specificity (M = 3.54, SD = 0.58), compared to other strategies." Additionally, "For Coordinated Collaboration, human-GPT agreement was highest (κ = 0.91), exceeding human-human agreement (κ = 0.83)." The framework successfully identified five behavioral codes, with four demonstrating "clear alignment with established theoretical frameworks" and one novel code (Peripheral Engagement) emerging with "high empirical consistency across coding batches."

Research paradigm

Mixed-methods (positivist/interpretivist integration through human-AI collaboration)

Author conclusions

The authors conclude that "GAIVA offers a practical and flexible approach for researchers who want to scale qualitative video analysis without losing interpretive depth. It is especially useful for studies focused on collaboration, embodiment, and classroom interaction-areas where rich multimodal data are essential but difficult to analyze manually at scale." They further state that "while GAIVA represents an important step toward scalable and rigorous qualitative video analysis, its continued evolution will depend on advancing multimodal capabilities, refining temporal granularity, improving robustness to real-world conditions, and developing community for prompt and theory transparency." Additionally, they note that "although GAIVA reduces manual labor, it does not eliminate subjective or theory-driven decisions that shape the scope and interpretation of findings."

Risk of bias

Selection bias: Only 13 of 24 students included in final analysis due to consent and data availability constraints; Small sample size: Five video recordings from single learning context (HoloOrbits simulation); Confirmation bias potential: Authors mitigated by instructing coders to evaluate whether claims were supported by observable evidence rather than searching for expected behaviors; Model bias: MLLMs trained on generic datasets may systematically underrepresent or mischaracterize culturally specific or contextually nuanced behaviors; Prompt bias: Researchers designed prompts that shape model outputs; could introduce theoretical preference bias; Coder drift: Although mitigated through high inter-rater reliability, subjective interpretations in codebook development could introduce bias; Selection bias: Only 13 of 24 students were included due to video consent and recording availability; Context-specificity: Training on general-purpose datasets (MLLMs/LLMs) may not align with context-sensitive learning behaviors; Confirmation bias: Mitigated through double-blind coding on initial subset with discussion-based resolution; Prompt engineering bias: Choice of theoretical frameworks and prompt design reflects researcher preferences; Sampling bias: Small sample size (5 video recordings, one hour each) limits generalizability; Model hallucination risk: MLLMs can produce plausible but unsupported statements about student intentions; Batch-processing effects: GPT-4o context window constraints required dividing transcriptions into 10 batches, potentially reducing exposure to less frequent behavioral patterns; Selection bias: Only 13 of 24 students included due to consent/data availability; Small sample size (5 video recordings, 13 students) limits generalizability; Model training data bias: MLLMs trained on general-purpose datasets may not represent educational contexts; Potential hallucination bias in MLLM outputs, noted as producing 'plausible but unsupported statements'; Batch processing bias: Context window limitations of GPT-4o may limit exposure to less frequent behavioral patterns; Researcher interpretation bias in codebook consolidation and theoretical alignment decisions; Confirmation bias potential in transcript evaluation (though partially mitigated by inter-rater reliability protocol)

Limitations

  • The authors state "the framework operates exclusively on vision-based data, without access to audio streams
  • As a result, it cannot capture verbal cues such as verbal discourse and audio-dependent collaborative behaviors such as speech patterns, vocal turn-taking, or prosodic indicators of engagement." Additionally, "the approach depends heavily on video quality and visibility requirements
  • Students must be clearly and consistently captured by cameras MLLMs to generate meaningful behavioral transcriptions
  • Situations involving occlusion, overlapping bodies, or dynamic camera movement may reduce the effectiveness of transcription and coding." The authors also note that "the use of fixed-length segmentation (e.g., 10-second clips) introduces artificial boundaries that may not align with the natural flow of interaction." Furthermore, "the current evaluation does not provide a claim-level error analysis or determine how specific transcription errors may propagate into downstream coding outcomes."

Open questions raised

  • Integration of audio data into MLLM transcription layer to capture verbal discourse and prosodic cues
  • Event-based segmentation approaches that identify natural behavioral boundaries rather than fixed temporal intervals
  • Training or fine-tuning MLLMs on diverse and noisy classroom video data to improve robustness
  • Long-context LLMs or retrieval-augmented summarization strategies to process larger amounts of video transcriptions
  • Claim-level error analysis to determine how specific transcription errors propagate into downstream coding outcomes
  • Tools for auditing model outputs and exposing analytic blind spots to improve transparency
Extracted from: pdfAgreement 54%

Explore related topics

Related papers