Interactive-Policy Distillation with Bidirectional Propose-and-Verify

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of teacher supervision caused by trajectory drift in online distillation by proposing an interactive policy distillation method. The core innovation lies in constructing a bidirectional propose-verify state machine that collaboratively generates hybrid trajectories and applies source-aware loss optimization, thereby unifying online and offline distillation paradigms to enable adaptive teacher intervention. From an engineering perspective, the approach integrates an inference engine, KV cache decoupling, and a state machine scheduling algorithm. Experimental results demonstrate that the proposed method improves mathematical reasoning accuracy by 3.28% and enhances data efficiency fourfold, significantly outperforming existing baselines.
📝 Abstract
On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
Teacher unanchoring
Knowledge distillation
Reasoning trajectory drift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interactive-Policy Distillation
Bidirectional Propose-and-Verify
Source-Split Loss
Fused Inference Engine
On-policy Distillation
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Shutong Wu
Shutong Wu
University of Wisconsin–Madison
Machine Learning
Xiwen Chen
Xiwen Chen
Clemson University
Deep LearningMultimodalComputer VisionTime Series AnalysisVLM/LLM
B
Brendan Rappazzo
Machine Learning Research, Morgan Stanley
D
Daiheng Zhang
Department of Electrical and Computer Engineering, Rutgers University–New Brunswick
Anderson Schneider
Anderson Schneider
Morgan Stanley
Machine Learning
Y
Yuriy Nevmyvaka
Machine Learning Research, Morgan Stanley
J
Jiawei Zhang
Department of Computer Sciences, University of Wisconsin–Madison