MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing multi-teacher knowledge distillation approaches, which rely on prompt-level domain labels and overlook cross-domain complementary signals, thereby constraining model generalization. To overcome these issues, this work proposes a label-free, token-level routing framework integrated with multi-teacher online policy distillation. Specifically, an expert alignment scoring mechanism is introduced to dynamically weight the supervisory signals from individual teachers, effectively mining and leveraging cross-domain complementary knowledge. By eliminating the dependency on domain labels, the proposed framework achieves substantial performance improvements over existing baselines in both labeled and unlabeled settings, yielding overall score gains of up to 12.3%.
📝 Abstract
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
Problem

Research questions and friction points this paper is trying to address.

Multi-teacher on-policy distillation
Teacher routing
Token-level routing
Unlabeled data
Cross-domain supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Token-level Routing
ExpertAlign
Unlabeled Data
Plug-in Interface
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
T
Tianze Xu
Shanghai Jiao Tong University
Y
Yanzhao Zheng
Alibaba Group
Z
Zhentao Zhang
Alibaba Group
Y
Yuanqiang Yu
Alibaba Group
C
Chao Ma
Alibaba Group
J
Jihuai Zhu
Alibaba Group
L
Lelun Wu
University of Science and Technology of China
Lyumanshan Ye
Lyumanshan Ye
Shanghai Jiao Tong Univeristy
Human-Computer Interaction
Pengfei Liu
Pengfei Liu
Associate professor at Shanghai Jiao Tong University
LLM
B
Baohua Dong
Alibaba Group
H
Hangcheng Zhu
Alibaba Group
R
Ruohui Huang
Alibaba Group
G
Gang Yu
Alibaba Group