From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
Weikang Shi, Houxing Ren, Junting Pan, Aojun Zhou, Ke Wang, Zimu Lu et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
The paper introduces KMP-Bench, a comprehensive K-8 Mathematical Pedagogical Benchmark with two complementary evaluation modules: KMP-Dialogue, which evaluates holistic pedagogical capabilities using multi-turn dialogue truncation and LLM-based evaluation with 22 criteria (4 general + 3 principle-specific for each of 6 pedagogical principles); and KMP-Skills, which assesses foundational tutoring abilities through multi-turn problem-solving, error detection/correction, and problem generation tasks.
Main result
The study reveals that "while current LLMs show promise in structured tasks like problem-solving, they struggle with the nuanced application of pedagogical principles and the reliable generation of pedagogically-sound responses." Additionally, "models fine-tuned on KMP-Pile dataset demonstrate substantial improvements across the KMP-Bench evaluations, underscoring the critical value of pedagogically-rich training data."
Research paradigm
empirical-positivist
Author conclusions
The authors conclude that "while LLMs have become exceptionally proficient at tasks with verifiable solutions, such as problem-solving and error correction, a significant performance gap remains in tasks requiring deep pedagogical awareness." They further state: "This distinction underscores that the frontier for AI tutors is shifting from being accurate problem-solvers to becoming effective, student-centric educators." Additionally, "Models fine-tuned on KMP-Pile dataset demonstrate substantial improvements across the KMP-Bench evaluations, underscoring the critical value of pedagogically-rich training data" for developing more effective AI math tutors.
Risk of bias
Evaluator bias: Using Gemini-2.0-Flash as primary evaluator may introduce systematic bias favoring certain model architectures or training paradigms; Selection bias: Seed problems drawn from nine specific datasets may not represent all K-8 mathematical content equally; Annotation bias: Although human evaluation shows 89.8% alignment with LLM evaluator, human annotators may have had implicit pedagogical preferences; Reference response bias: Using original tutor turns as reference responses assumes quality and appropriateness of those original responses; Selection bias in seed problem sources (9 datasets may not represent all K-8 math content equally); Evaluator bias: LLM-based evaluation may inherit biases from Gemini-2.0-Flash model; Human annotation bias in the manual verification steps (451 flows removed at 7.6% rate); Training data bias: KMP-Pile contains 150K dialogues generated by Gemini-2.0-Flash, which may introduce systematic biases; Potential funding bias from affiliation with Multimedia Laboratory and institutional resources; Potential bias in LLM-based evaluation: reliance on Gemini-2.0-Flash as evaluator may introduce systematic biases; Human annotation bias: only 300 instances manually evaluated for validation (limited sample); Selection bias in seed problems: problems sourced from nine datasets may not represent all educational contexts; Temporal bias: evaluation conducted at specific time point with models available in 2026; Evaluator alignment: 89.8% agreement between LLM evaluator and human judgment suggests 10.2% disagreement rate
Limitations
- The authors state that "The scope of our benchmark is currently confined to the K-8 curriculum, excluding higher-level mathematics, and future work could extend this coverage to assess and enhance tutoring capabilities for more advanced subjects
- Additionally, our framework is entirely text-based, while effective mathematics education often incorporates visual elements like diagrams and graphs
- future work could explore multi-modal interactions to create a more holistic tutoring experience
- Furthermore, we only employed supervised fine-tuning to enhance the performance of open-source LLMs, without utilizing other post-training strategies such as Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO)."
Open questions raised
- Extension of benchmark scope beyond K-8 curriculum to higher-level mathematics
- Multi-modal interactions incorporating visual elements (diagrams, graphs) for mathematics education
- Application of advanced post-training strategies (RLHF, DPO) beyond supervised fine-tuning
- Assessment of pedagogical reasoning and reliable content generation in LLMs
- Current evaluations of LLMs in mathematical tutoring rely on simplistic metrics (problem-solving accuracy, textual similarity scores like BLEU, BERTScore) that inadequately reflect teaching efficacy in dynamic, real-world educational settings
- Existing assessments are confined to narrow scenarios (e.g., error correction only) and overlook critical tutoring functions such as proactive follow-up questioning, student confusion clarification, and guided problem generation
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations
- Revolutionizing education with AI: Exploring the transformative potential of ChatGPTTufan Adıgüzel · 2023 · 858 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations