🤖 AI Summary
This study addresses the capacity gap in reasoning distillation arising from distributional discrepancies between teacher and student models, which hinders conventional methods from simultaneously achieving high supervision quality and distillation efficiency. To overcome this limitation, this work proposes TeacherGRPO, a framework that directly adapts the teacher model to the student's distribution via reinforcement learning. Building upon Group Relative Policy Optimization (GRPO), the approach introduces a curriculum-based selective alignment mechanism to focus on high-signal reasoning discrepancies. Furthermore, an importance-adaptive length regularization is designed to suppress redundancy while preserving critical reasoning steps. Extensive evaluations demonstrate that TeacherGRPO significantly outperforms existing baselines across multiple reasoning benchmarks, effectively enhancing the reasoning capabilities of smaller models. The implementation code has been made publicly available.
📝 Abstract
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.