🤖 AI Summary
This study addresses the limitations of large language models in multi-turn pedagogical interactions, specifically their inability to accumulate and reuse instructional experience or adapt to diverse student needs. To this end, we propose the REAT framework, which introduces a novel Observer-Critic-Mentor multi-agent distillation pipeline that structures teaching experiences from historical dialogues into transferable knowledge capable of generalizing across models. Furthermore, by integrating state-aware retrieval techniques, REAT enables real-time experience invocation and adaptive tutoring. Experimental results demonstrate that REAT significantly outperforms prompt engineering and supervised fine-tuning baselines on mathematical tutoring tasks, exhibiting particularly strong performance in complex, low-proficiency scenarios.
📝 Abstract
Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.