on-policy distillation

Design and implement training algorithms, loss functions, and data pipelines that transfer or compress behavior policies by learning from trajectories produced by the student policy itself (on-policy), producing a deployable student policy aligned to teacher signals, outcome calibration, or desired behavioral constraints. This includes methods for self- and co-distillation, bidirectional or annealed update schedules, forward‑KL or logit‑free losses, reward‑gated or selective supervision, sparse/visual/annotation‑free variants, on‑policy sampling adaptation and calibration, and mechanisms to arbitrate heterogeneous teacher signals or retain prior platform behaviors during continual updates.

on-policydistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$233K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This work systematically investigates on-policy distillation (OPD) for large language models to address the exposure bias arising from train-test mismatch in conventional off-policy knowledge distillation. We introduce, for the first time, a unified f-divergence theoretical framework that categorizes and integrates existing techniques along three orthogonal dimensions: feedback signal, teacher access mode, and loss granularity—encompassing white-box, black-box, and teacher-free settings as well as token-level and sequence-level losses. The study reveals an intrinsic connection between OPD and interactive imitation learning, reviews representative methods and industrial practices, and identifies key open challenges such as distillation scaling laws and uncertainty-aware feedback, thereby providing a clear technical roadmap for future research.

Exposure BiasImitation LearningKnowledge Distillation

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenges of sparse rewards in reinforcement learning and performance degradation caused by teacher-student capability mismatch during self-distillation. To this end, we propose JOLT, a method that jointly trains a single policy to serve as both a privileged teacher and an unprivileged student, ensuring that the guidance remains aligned with the student’s current capabilities. By deriving the necessary and sufficient conditions under which the teacher update constitutes a positive multiple of the student gradient, we design a teacher optimization objective that integrates outcome rewards with KL regularization. Leveraging on-policy distillation and joint optimization, JOLT significantly improves both training efficiency and final performance on tasks such as mathematical reasoning and programming.

On-Policy DistillationPrivileged InformationReinforcement Learning

This study addresses the performance degradation of teacher models in online knowledge distillation caused by processing off-policy prefixes generated by student models. To tackle this issue, we propose SCOUT, a framework that systematically mitigates off-policy bias for the first time. Through co-training, the teacher adapts to student-generated trajectories, while reinforcement learning with verifiable rewards drives adaptive teacher updates to optimize its conditional continuation capability. Our approach significantly enhances the quality of teacher continuations given student prefixes and consistently improves distillation effectiveness across diverse configurations and reasoning tasks.

Off-Policy TeacherOn-Policy DistillationPrefix Degradation

This study addresses the challenge in online policy distillation where the continuous evolution of student policies causes early trajectory divergences to be overlooked, rendering fixed sampling or replay strategies suboptimal for supervision budget utilization. To overcome this, we propose R-OPD, a curriculum learning framework that introduces a novel adaptive scheduling strategy based on a gradient-triggering mechanism. By performing minibatch-level variance detection to monitor gradient dynamics, our method precisely identifies the onset of policy degradation and dynamically switches to initial policy replay, thereby facilitating efficient knowledge transfer. Extensive experiments demonstrate that R-OPD significantly improves accuracy across multiple mathematical reasoning benchmarks, achieving gains of up to 6.15%, while generating longer reasoning chains that fully unlock the model's potential.

Curriculum LearningKnowledge DistillationOn-Policy Distillation

Standard online policy distillation (OPD) often suffers from training instability due to high noise in teacher trajectories and large variance in supervision signals. To address this, this work proposes the BRTS framework, which introduces a novel multi-trajectory sampling and prioritization mechanism that selects high-quality teacher trajectories primarily based on their correctness and secondarily on their alignment with the student’s behavior. Additionally, BRTS incorporates a ground-truth conditioned recovery strategy to handle challenging samples and integrates an auxiliary supervision loss to enhance training stability. Evaluated on demanding mathematical reasoning benchmarks—including AIME 2024/2025 and AMC 2023—BRTS significantly outperforms standard OPD, with the most pronounced gains observed on the hardest problem subsets.

High-VarianceOn-Policy DistillationReasoning

Reinforcement Teaching

Apr 25, 2022
AL
Alex Lewandowski
🏛️ University of Alberta | Huawei Technologies Canada Co., Ltd. | Google Brain

Existing meta-learning methods suffer from limited generalizability, often being confined to specific algorithms or requiring differentiability assumptions. This paper proposes a general reinforcement learning–driven meta-learning framework that trains a teacher policy to dynamically guide arbitrary student algorithms—without imposing structural or differentiability constraints on the student. Key contributions include: (i) the first unified pedagogical paradigm for meta-learning; (ii) a parameter-behavior encoder that implicitly infers the student’s internal parameter state from its input-output behavior; and (iii) a reward function grounded in learning progress. Experiments across supervised and reinforcement learning tasks demonstrate that our framework significantly outperforms baselines relying on heuristic rewards and handcrafted state representations, validating its broad generalizability and empirical effectiveness.

AdaptabilityMachine Learning EfficiencyMeta-Learning

Latest Papers

What's happening recently
View more

This study addresses the high computational cost of teacher supervision in online knowledge distillation by proposing an adaptive supervision selection mechanism that leverages the student model's own successful trajectories as a reference. Specifically, this method precisely allocates teacher resources by filtering out failed samples, while integrating hidden state trajectory discrepancy analysis with a budget-aware token selection strategy to achieve efficient supervision. Experimental results demonstrate that the proposed approach maintains comparable inference performance while requiring only 3.46% to 5.02% of the teacher inputs, thereby substantially reducing the computational overhead associated with online distillation.

Knowledge distillationOn-policy distillationTeacher computation cost

This study addresses the challenge in multi-turn agent online distillation where early student errors accumulate, causing trajectory deviation from the teacher distribution and rendering supervision ineffective. To overcome this, we propose the STI-OPD framework, which introduces a novel stochastic intervention mechanism based on policy divergence to replace fixed thresholds, dynamically substituting student actions to balance control with exploration. Furthermore, an importance-weighted reverse KL loss is incorporated to correct sampling bias in mixed trajectories, thereby enhancing optimization reliability. Comprehensive evaluations on tool-use reasoning and long-horizon interaction benchmarks demonstrate that our approach consistently outperforms state-of-the-art baselines, validating the effectiveness of both the stochastic intervention strategy and the importance weighting module.

Distribution driftError accumulationMulti-turn agentic tasks

This study addresses the cold-start problem caused by sparse rewards in reinforcement learning for long-horizon agents by proposing the GATS method. GATS employs online policy distillation to provide token-level guidance signals and introduces a novel finding that distillation benefits depend on the performance gap between teacher and student models. Based on this insight, an adaptive weight scheduling mechanism is designed to automatically attenuate and eventually remove teacher guidance as the gap narrows, enabling a smooth transition from guided learning to autonomous exploration while supporting teacher models smaller than the student. Evaluated on benchmarks such as ALFWorld, GATS improves success rates by 4.37%–11.87% over GRPO and achieves the best average performance across all configurations.

Cold-Start ProblemLong-Horizon AgentsOn-Policy Distillation

This study addresses the challenge in multi-teacher knowledge distillation where base model preferences interfere with post-training signals, thereby constraining policy composition and routing performance. To overcome this, we propose the Δ-MOPD framework, which introduces a novel objective construction mechanism based on relative teacher shifts. By transferring log-probability shifts and anchoring student initialization, this approach effectively decouples base preferences from post-training increments, establishing objective construction as an independent design dimension. Experimental results demonstrate that under a three-teacher configuration, the method yields a 4.11-point improvement on mathematical tasks and an average gain of 1.95 points across five benchmarks. Furthermore, staged routing significantly reduces sequential discrepancies to 6.42 points. These findings validate the efficacy of the proposed framework for multi-teacher online distillation.

endpoint policyknowledge transferlogit shift

This study addresses the misalignment between supervision signals and final rewards caused by local distribution shifts in multi-turn agent policy distillation. To overcome this limitation, we propose Outcome-Guided Online Policy Distillation (OG-OPD), which transcends conventional local optimization by introducing a dynamic supervision mechanism based on relative trajectory weights. Specifically, OG-OPD calibrates teacher guidance signals using ultimate task outcomes, enabling precise intervention at critical steps essential for improving success rates. Experimental results demonstrate that OG-OPD significantly outperforms existing baselines on benchmarks such as ALFWorld, achieving an improvement in task success rate of up to 17.7 percentage points.

multi-turn autonomous agentson-policy distillationoutcome-guided supervision

Hot Scholars

LB

Lei Bai

Shanghai AI Laboratory
Foundation ModelScience IntelligenceMulti-Agent SystemAutonomous Discovery
GS

Guanya Shi

Assistant Professor, CMU RI | Amazon Scholar, FAR (Frontier AI & Robotics)
RoboticsRobot LearningReinforcement LearningControl
JL

Jian Luan

Toshiba, Microsoft, Xiaomi
LLMVLMTTSSinging Synthesis
PA

Pieter Abbeel

UC Berkeley | Covariant
RoboticsMachine LearningAI