Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
This study addresses the limitation of existing multi-teacher online distillation methods, which exploit only output distributions while neglecting internal representations. To overcome this, we propose the first representation-level multi-teacher online distillation framework. Specifically, this work introduces a novel dual-channel mechanism that integrates hidden states with token predictions to coordinate knowledge from multiple experts without additional training. Furthermore, it incorporates late-layer target selection, shared projection bridging, and domain-grouped updating techniques to enable dynamic adjustment of supervision weights. Extensive evaluations across nine benchmarks, including mathematical reasoning and code generation, demonstrate that the student model consistently outperforms baseline approaches and surpasses most single-domain expert teachers in overall performance.