π€ AI Summary
This study addresses the inter-domain conflicts in multi-teacher online distillation caused by the momentum and adaptive scaling of optimizers such as AdamW, which render gradient correction ineffective. To overcome this, we propose an update projection method that, for the first time, shifts projection constraints from the gradient space to the optimizer update space. Specifically, after generating candidate displacements from mixed gradients, projections are applied exclusively to those violating the constraints, ensuring parameter updates do not increase the loss in any domain. This approach bridges a critical theoretical gap in gradient correction under complex optimizers. Experiments demonstrate that our method improves IFEval by 2.96 points (averaging 60.03) on medical and general tasks, while achieving state-of-the-art average performance of 32.67 on mathematics and code benchmarks.
π Abstract
On-policy distillation from multiple teachers combines expertise from different domains in a single student, but conflicting gradients can hinder this integration. Gradient corrections directly constrain parameter updates under plain SGD. With optimizers such as AdamW, however, momentum, adaptive scaling, and weight decay can turn a corrected gradient into an update that increases a domain loss to first order. To address this gap, we propose Update Projection for Multi-Teacher On-Policy Distillation (UP-MOPD). UP-MOPD lets the original mixed gradient update the optimizer state and generate a candidate displacement, then projects only violating candidates before they are committed to the parameters. The projection gives the unique feasible update closest to the candidate in Euclidean distance. In experiments combining medical and general domains, UP-MOPD improves IFEval-loose accuracy late in training by 2.96 points over vanilla M-OPD. It achieves an average score of 60.03 across eight metrics, compared with 59.00 for gradient projection and 59.15 for update rejection. On a public benchmark covering mathematics, code, and instruction following, it achieves the best average across six tasks (32.67), leads on LiveCodeBench v5, and ties for the best IFEval result.These results support projecting optimizer updates to reduce interference between domains.