UP-MOPD: Update Projection in Multi-Teacher On-Policy Distillation

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inter-domain conflicts in multi-teacher online distillation caused by the momentum and adaptive scaling of optimizers such as AdamW, which render gradient correction ineffective. To overcome this, we propose an update projection method that, for the first time, shifts projection constraints from the gradient space to the optimizer update space. Specifically, after generating candidate displacements from mixed gradients, projections are applied exclusively to those violating the constraints, ensuring parameter updates do not increase the loss in any domain. This approach bridges a critical theoretical gap in gradient correction under complex optimizers. Experiments demonstrate that our method improves IFEval by 2.96 points (averaging 60.03) on medical and general tasks, while achieving state-of-the-art average performance of 32.67 on mathematics and code benchmarks.
πŸ“ Abstract
On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval result.These results support projecting optimizer updates to reduce interference between domains.
Problem

Research questions and friction points this paper is trying to address.

multi-teacher distillation
on-policy distillation
gradient conflict
optimizer update interference
domain loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Update Projection
Gradient Conflict
AdamW Optimizer
Domain Interference
πŸ”Ž Similar Papers
T
Taojie Zhu
Tsinghua University
J
Jing Jin
Tsinghua University
Y
Yuan Xia
Ant Group
C
Chenyang Ding
Tsinghua University
Q
Qunshan He
Zhejiang University
W
Wanke Xia
Tsinghua University
Tao Sun
Tao Sun
Ant Group
Y
Yan Chen
Tsinghua University
Jian Wang
Jian Wang
Senior Staff Algorithm Engineer, Ant Group
Computer VisionMultimodalLLM
Jinjie Gu
Jinjie Gu
ant group
ζœΊε™¨ε­¦δΉ οΌŒζŽ¨θ
T
Tao Feng
Tsinghua University