Sherpa: Teaching LLMs to Teach Adaptively

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing large language model (LLM) tutoring systems that rely on static criteria, which fail to accommodate individual student differences. We propose a multi-turn reinforcement learning framework that utilizes authentic learning outcomes as reward signals. By simulating diverse student archetypes and incorporating preference-conditioned modeling, the teacher model is trained to directly optimize learning efficacy, thereby enabling adaptive, personalized instruction. Experimental results demonstrate that this approach improves student performance by 20.5%, achieves a pedagogical quality score of 79.2%, and attains a human preference rate of 79.6%. This work transcends conventional static paradigms, offering a novel pathway for LLM-driven personalized education.
📝 Abstract
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Adaptive Teaching
AI Tutors
Student Learning Outcomes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-turn Reinforcement Learning
Student Archetypes
Adaptive Teaching
Large Language Models
AI Tutors