From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear mechanisms by which teacher signals influence parameter updates and capability integration in multi-teacher online distillation. Using models such as Qwen3, we conduct reinforcement learning and multi-teacher distillation experiments, comparing gradients, optimizer updates, and learning curves. We provide the first quantitative analysis of implicit weight changes induced by Adam’s first-moment smoothing and BF16 rounding, revealing the underlying mechanisms of loss averaging strategies. Our findings demonstrate that response-level averaging improves mathematical accuracy by 2.6 percentage points, whereas global token-level averaging leads to performance degradation. By elucidating the differential impacts of distinct averaging strategies on task performance, this work offers a novel perspective for understanding distillation sensitivity.
📝 Abstract
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Problem

Research questions and friction points this paper is trying to address.

multi-teacher distillation
on-policy distillation
gradient analysis
knowledge distillation
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Gradient Analysis
Optimizer Dynamics
Low-Precision Quantization
Knowledge Distillation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Siqi Zhu
University of Illinois Urbana-Champaign
S
Suozhi Huang
Princeton University
K
Kaixuan Zhang
University of Illinois Urbana-Champaign
Y
Yuheng Yang
Westlake University
Z
Zhanyang Jin
University of Illinois Urbana-Champaign
Y
Yihang Sun
University of Illinois Urbana-Champaign
Jiaxuan You
Jiaxuan You
Assistant Professor, UIUC CS
Foundation ModelsGNNLarge Language Models