Score
Designs and implements procedures that compress an expert or privileged ‘teacher’ policy into a deployable ‘student’ policy by aggregating teacher demonstrations and corrective labels using the DAgger (dataset aggregation) protocol and applying supervised distillation; may include handling privileged teacher inputs at training time and optionally fine‑tuning the distilled student with reinforcement learning to recover or improve closed‑loop performance.
This study systematically investigates the feedback-to-update mechanism in on-policy distillation (OPD), where data are generated by the current policy. Framing OPD as a feedback-to-update problem, the work introduces a formula-driven categorization framework that unifies two major update pathways: distributional loss and policy gradient–style log-ratio updates. It further incorporates novel perspectives from temporal credit assignment and temporal vocabulary routing. By leveraging techniques such as KL divergence orientation, generalized advantage estimation (GAE), and counterfactual routing, the analysis reveals that OPD performance critically depends on state compatibility and support set construction. The paper establishes a comprehensive analytical framework for OPD, derives explicit bias bounds, and proposes new methods—GAE-OPD and CR-OPD—to enhance training stability, alongside actionable diagnostic tools and a practical implementation checklist.
This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.
This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.
This paper addresses policy distillation under privileged information: the teacher observes the full state, whereas the student accesses only partial observations—inducing information asymmetry, distributional shift, and policy degradation. Existing approaches either degrade teacher capability to generate realizable demonstrations or compel the student to blindly explore unobserved states, both yielding low sample efficiency. We propose an active querying–correction mechanism and intelligent reinitialization to construct recoverable trajectories within the student’s observable subspace, avoiding forced imitation of unrealizable teacher policies. Our method integrates adaptive-query imitation learning with recovery-state–based reinforcement learning, requiring no teacher modification or auxiliary exploration. Evaluated on both simulation and real-robot tasks, our approach significantly improves training efficiency and final performance, consistently outperforming standard teacher–student distillation baselines.
This study addresses the challenge in online policy distillation where teacher intervention alters the student's distribution, making it difficult for fixed-intensity strategies to balance trajectory quality against off-policy deviation. To overcome this limitation, this work proposes MAESTRO, a framework that adaptively regulates both the timing of teacher takeover and the generation duration through a segment-aggregated policy divergence scoring mechanism, thereby achieving joint dynamic adaptation of intervention depth and length. Empirical evaluations demonstrate that MAESTRO attains state-of-the-art average accuracy across eight mathematical reasoning benchmarks. Furthermore, it reduces training response lengths by 67.3% compared to standard methods, substantially improving knowledge transfer efficiency.
This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.
This study addresses the performance degradation of teacher models in online knowledge distillation caused by processing off-policy prefixes generated by student models. To tackle this issue, we propose SCOUT, a framework that systematically mitigates off-policy bias for the first time. Through co-training, the teacher adapts to student-generated trajectories, while reinforcement learning with verifiable rewards drives adaptive teacher updates to optimize its conditional continuation capability. Our approach significantly enhances the quality of teacher continuations given student prefixes and consistently improves distillation effectiveness across diverse configurations and reasoning tasks.
This study addresses the feedback scale imbalance in multi-teacher online policy distillation (MOPD), which hinders student models from uniformly assimilating capabilities across diverse domain experts. To overcome this, we propose a domain-normalized MOPD framework that, for the first time, reveals and quantifies feedback variance bias within MOPD. By introducing an adaptive weighting mechanism grounded in empirical distributions to dynamically rescale feedback intensities across domains, our approach transcends the limitations of conventional fixed routing strategies. Integrating large language models with reinforcement learning techniques, extensive experiments on the Qwen3.5 model series demonstrate that the proposed method consistently outperforms baselines across six public benchmarks. Notably, it significantly recovers performance degradation in mathematical reasoning while achieving efficient multi-skill integration.
This study addresses the sensitivity of performance to learning rates in federated online distillation, which arises from the coupling between aggregation and generated feedback. Specifically, it reveals an optimization lag mechanism induced by the dual role of student models. To mitigate this issue, this work proposes FedTOPS, a method that theoretically analyzes the aforementioned coupling effect and designs an adaptive scaling strategy based on predictive change constraints to enable dynamic updates. By integrating federated learning with online knowledge distillation, FedTOPS substantially enhances multi-model collaborative training. Experimental results across six benchmarks demonstrate that the proposed approach improves macro-average scores by 4.56 to 14.57 percentage points over FedAvg, confirming its effectiveness in optimizing federated online distillation frameworks.
This study addresses the challenge in mixed-strategy distillation where teachers' inherent preferences interfere with student models and feedback guidance is underutilized. To overcome these limitations, this work proposes an inductive feedback method that introduces a probabilistic confirmation framework to decouple teacher preferences from feedback signals for constructing target distributions. Furthermore, it designs a shared rolling estimator based on symmetric divergence minimization, which integrates importance weighting with trust region constraints to maximize feedback utilization efficiency. Extensive evaluations demonstrate that the proposed approach significantly outperforms conventional online distillation and contrastive variants across both knowledge and agent benchmarks, effectively enhancing the post-training performance of language models.
This study addresses the failure of teacher supervision caused by trajectory drift in online distillation by proposing an interactive policy distillation method. The core innovation lies in constructing a bidirectional propose-verify state machine that collaboratively generates hybrid trajectories and applies source-aware loss optimization, thereby unifying online and offline distillation paradigms to enable adaptive teacher intervention. From an engineering perspective, the approach integrates an inference engine, KV cache decoupling, and a state machine scheduling algorithm. Experimental results demonstrate that the proposed method improves mathematical reasoning accuracy by 3.28% and enhances data efficiency fourfold, significantly outperforming existing baselines.
This study addresses the challenge in multi-teacher knowledge distillation where base model preferences interfere with post-training signals, thereby constraining policy composition and routing performance. To overcome this, we propose the Δ-MOPD framework, which introduces a novel objective construction mechanism based on relative teacher shifts. By transferring log-probability shifts and anchoring student initialization, this approach effectively decouples base preferences from post-training increments, establishing objective construction as an independent design dimension. Experimental results demonstrate that under a three-teacher configuration, the method yields a 4.11-point improvement on mathematical tasks and an average gain of 1.95 points across five benchmarks. Furthermore, staged routing significantly reduces sequential discrepancies to 6.42 points. These findings validate the efficacy of the proposed framework for multi-teacher online distillation.