🤖 AI Summary
This study addresses the limitation of existing multi-teacher online distillation methods, which exploit only output distributions while neglecting internal representations. To overcome this, we propose the first representation-level multi-teacher online distillation framework. Specifically, this work introduces a novel dual-channel mechanism that integrates hidden states with token predictions to coordinate knowledge from multiple experts without additional training. Furthermore, it incorporates late-layer target selection, shared projection bridging, and domain-grouped updating techniques to enable dynamic adjustment of supervision weights. Extensive evaluations across nine benchmarks, including mathematical reasoning and code generation, demonstrate that the student model consistently outperforms baseline approaches and surpasses most single-domain expert teachers in overall performance.
📝 Abstract
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.