🤖 AI Summary
This study addresses the unclear mechanisms by which teacher signals influence parameter updates and capability integration in multi-teacher online distillation. Using models such as Qwen3, we conduct reinforcement learning and multi-teacher distillation experiments, comparing gradients, optimizer updates, and learning curves. We provide the first quantitative analysis of implicit weight changes induced by Adam’s first-moment smoothing and BF16 rounding, revealing the underlying mechanisms of loss averaging strategies. Our findings demonstrate that response-level averaging improves mathematical accuracy by 2.6 percentage points, whereas global token-level averaging leads to performance degradation. By elucidating the differential impacts of distinct averaging strategies on task performance, this work offers a novel perspective for understanding distillation sensitivity.
📝 Abstract
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.