🤖 AI Summary
This study questions whether the performance gains of multi-teacher online distillation (MOPD) stem from the algorithm itself or from hyperparameter optimization. Through controlled experiments across multiple models and benchmarks, we systematically compare supervised fine-tuning, soft-label distillation, mixed-teacher prefix distillation, and weight merging. Our results demonstrate that thoroughly tuned offline methods achieve accuracy nearly identical to MOPD, yet the latter incurs 15 to 23 times higher computational overhead. This work reveals a stark mismatch between the prohibitive cost and marginal benefits of MOPD, challenging prevailing consensus in the field. Furthermore, it establishes that carefully optimized offline baselines serve as a substantially more cost-effective alternative for knowledge transfer in language models.
📝 Abstract
Combining capabilities of multiple expert models trained starting from the same base checkpoint has become increasingly common in frontier language-model post-training. Recent trends suggest that multi-teacher on-policy distillation (MOPD) outperforms conventional off-policy methods. However, despite the higher inference and environment interaction costs incurred by MOPD, we find that much of its reported accuracy gain is due to certain training design choices and hyperparameter optimization, rather than the algorithm itself. We conduct a controlled self-distillation study across two multi-teacher settings, four models, and eleven benchmarks, comparing off-policy methods, namely supervised fine-tuning (SFT) and soft-label distillation, to hybrid teacher-prefix distillation and MOPD. We found that all four methods achieve \textit{nearly identical} accuracy. However, MOPD uses $14.8$--$23.1\times$ SFT's training GPU-hours. We also revisit four recently published comparisons between on-policy and off-policy distillation and find that the reported on-policy gains shrink substantially when SFT baselines are trained on rejection-sampled teacher trajectories and use independently tuned hyperparameters. As a training-free alternative, we also find that simple weight-merging methods can recover expert capabilities with minutes of CPU merging time, although their accuracy degrades as model size decreases and task interference increases. Overall, our results question recent gains reported due to MOPD and suggest careful tuning of more efficient off-policy baselines as a viable alternative.