Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the feedback scale imbalance in multi-teacher online policy distillation (MOPD), which hinders student models from uniformly assimilating capabilities across diverse domain experts. To overcome this, we propose a domain-normalized MOPD framework that, for the first time, reveals and quantifies feedback variance bias within MOPD. By introducing an adaptive weighting mechanism grounded in empirical distributions to dynamically rescale feedback intensities across domains, our approach transcends the limitations of conventional fixed routing strategies. Integrating large language models with reinforcement learning techniques, extensive experiments on the Qwen3.5 model series demonstrate that the proposed method consistently outperforms baselines across six public benchmarks. Notably, it significantly recovers performance degradation in mathematical reasoning while achieving efficient multi-skill integration.
📝 Abstract
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Problem

Research questions and friction points this paper is trying to address.

Multi-teacher distillation
Reinforcement learning
Feedback imbalance
Language model merging
On-policy distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-teacher on-policy distillation
Domain normalization
Feedback rescaling
Reinforcement learning
Knowledge distillation
🔎 Similar Papers
No similar papers found.
X
Xin Li
Nanyang Technological University
H
Hao Jiang
Nanyang Technological University
X
Xin Gao
Yale University
Annan Wang
Annan Wang
PhD at NTU CCDS
Computer Vision and VLM
Y
Yuchen Xie
Nanyang Technological University
J
Jinghao Guo
Nanyang Technological University
X
Xingwei Qu
University of Manchester
Y
Yichi Zhang
Chau Yuen
Chau Yuen
IEEE Fellow, Highly Cited Researcher, Nanyang Technological University
WirelessSmart GridLocalizationIoTBig Data