12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Benchmark-based evaluation study involving: (1) construction of a risk taxonomy grounded in learning sciences literature; (2) curation of seed questions from established datasets; (3) systematic generation of single-turn (n=3,135) and multi-turn (n=2,820) pedagogically unsafe interactions using specialized prompting and crescendo-based escalation methodology; (4) human validation through three complementary annotation stages (domain validity, risk alignment, crowd generalizability); (5) evaluation of 11 LLMs (10 open-weight, 1 closed-weight, 3.8B–72B parameters) using automated harm scoring via majority voting across three evaluator models and human evaluation by doctoral students with STEM teaching experience..

Sample

N = 11, 6 groups

Primary method

Harm rate calculation: HR = (Number of unsafe outputs)/(Total outputs generated), reported as percentage. Agreement metrics: Fleiss' κ for multi-rater agreement across three annotation stages. Final label determination: majority voting across three evaluator LLMs (GPT-5.2, DeepSeek-32B, Claude Sonnet 4.5) for automated harm scoring; three-way majority vote or discussion-based resolution for human annotations. Stratified sampling used for human evaluation sample selection. No inferential statistical tests (t-tests, ANOVAs, chi-squares) are reported.

Main result

The study found that "no model is reliably safe - every model exceeds 60% harm rate on at least five dimensions in single-turn and six in multi-turn." Additionally, "multi-turn interaction amplifies rather than corrects harm, average harm increases by 6–11 percentage points across subjects, and Pedagogical harm undergoes the largest shift in the benchmark, surging from a cross-model average of 17.7% in single-turn to 77.8% in multi-turn." The results also show that "larger scale does not consistently improve safety within the Qwen2.5 family, scaling from 7B to 72B yields improvement on some dimensions but regression on others."

Reports effect sizes.

Research paradigm

Empirical-positivist with pragmatist educational grounding

Author conclusions

The authors conclude that "models which appear helpful or safe in isolated responses can fail as tutors over extended dialogue, underscoring the need for evaluation that jointly measures learning support and harm avoidance." They further state that "harm profiles are subject-dependent - mathematics shows the highest metacognitive harm (92.8%) but the lowest epistemic harm (26.0%) in multi-turn, demonstrating that mitigation strategies must be discipline-aware." The work demonstrates that "neither parameter count nor proprietary alignment alone yields reliable safety improvements; targeted pedagogical alignment is necessary."

Risk of bias

Selection bias in seed question curation: Questions were sampled from MathDial, CAMEL-AI chemistry, and CAMEL-AI physics datasets, potentially over-representing certain problem types or domains.; Annotator expertise bias: Stage 2 (risk alignment) restricted to 3 doctoral students with teaching experience; potential convergence toward academic rather than practitioner-grounded perspectives.; Evaluator model bias: Harm assessment relies on three LLM evaluators (GPT-5.2, DeepSeek-32B, Claude Sonnet 4.5); their own biases and pedagogical assumptions may propagate to harm scoring.; Synthetic dialogue bias: Multi-turn conversations generated via crescendo methodology may not reflect authentic student behaviors and escalation patterns.; Domain specificity: Evaluation restricted to STEM subjects (physics, chemistry, mathematics); findings may not generalize to humanities or social sciences.; Crowdworker bias: Stage 3 annotation by 24 Prolific workers filtered for STEM undergraduate degree; may underrepresent non-technical perspectives on pedagogical harm.; Annotation bias: Decreasing inter-rater agreement across annotation stages (0.82 to 0.69) suggests increasing subjectivity, particularly for pedagogical judgments.; Selection bias: Seed questions drawn from established datasets (MathDial, CAMEL-AI) may not represent full spectrum of student queries.; Evaluator model bias: Harm scoring relies on majority voting among three LLMs (GPT-5.2, DeepSeek-32B, Claude Sonnet 4.5), introducing potential biases from these evaluator models themselves.; Domain limitation: Study restricted to three STEM subjects; generalizability to humanities or other domains unknown.; Synthetic dialogue bias: Dataset comprises generated (not real) student-tutor interactions, potentially missing authentic student behavior patterns.; Human evaluation sample bias: Only 1,500–2,400 turns evaluated by humans across all models and subjects (small sample relative to full evaluation set).; Selection bias in seed question curation from existing datasets (MathDial, CAMEL-AI) may not represent full diversity of student queries; Synthetic generation of risky interactions via crescendo-based methodology may not capture authentic tutoring failure modes; Automated evaluation using LLM judges introduces circularity and potential alignment bias from the evaluator models themselves; Limited human evaluation sample (30% of benchmark) with only 2-3 annotators per instance may not capture full ground truth variance; Annotation agreement decreases with expertise level (κ: 0.82→0.74→0.69 across stages), suggesting subjective judgment difficulty; Majority voting resolution may mask systematic disagreement on pedagogically subjective judgments; Imbalanced model coverage: 10 open-weight models vs. 1 proprietary model (GPT-5-mini)

Open questions raised

  • Current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture whether a model is simultaneously pedagogically effective and safe across student-tutor interaction.
  • Existing tutoring benchmarks (MathDial, MathTutorBench) evaluate pedagogical quality but do not systematically characterize tutoring-specific safety risks.
  • Safety benchmarks (RealToxicityPrompts, DecodingTrust, SG-Bench, CoSafe, CASEBench) probe toxicity and adversarial robustness but are not grounded in educational objectives or student misconceptions.
  • Both strands of work rely predominantly on single-turn interaction; multi-turn trajectories where pedagogical and safety failures accumulate are understudied in educational settings.
  • No benchmark jointly evaluates adversarial robustness and instructional quality in educational settings with multi-turn dialogue.
  • Generative tutors require new instrumentation that jointly addresses robustness and pedagogical quality; existing ITS detection methods target constrained output spaces.
Data: SAFETUTORS benchmark (code & dataset): https://github.com/RadiantCrystal/SafeTutors; MathDial (seed questions for mathematics): Macina et al., 2023; CAMEL-AI Chemistry dataset: https://huggingface.co/datasets/camel-ai/chemistry; CAMEL-AI Physics dataset: https://huggingface.co/datasets/camel-ai/physics; Yes, the paper states: "Code & Dataset https://github.com/RadiantCrystal/SafeTutors". Seed questions sourced from MathDial (Macina et al., 2023) and CAMEL-AI chemistry and physics datasets.; SAFETUTORS benchmark dataset (3,135 single-turn instances and 2,820 multi-turn conversations) available at https://github.com/RadiantCrystal/SafeTutorsCode: https://github.com/RadiantCrystal/SafeTutors; https://github.com/RadiantCrystal/SafeTutors (code and dataset)Extracted from: pdfAgreement 42%

Explore related topics

Related papers