A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.