PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the capability seesaw effect and cross-task interference caused by shared parameters in multi-teacher online distillation. To mitigate these issues, we propose a geometry-aware subspace protection mechanism that constructs subspace memory by analyzing the geometric properties of parameter updates and projects gradients to eliminate interference. Furthermore, a lightweight conflict probe is designed to optimize task ordering, complemented by a cyclic scheduling strategy that balances estimation accuracy with task revisit frequency, thereby achieving equitable integration of multiple capabilities. Experimental results demonstrate that the proposed method consistently outperforms baselines across code, reasoning, and mathematics tasks, yielding average improvements of 2.54 and 2.09 points on Qwen2.5-7B and Llama-3.1-8B, respectively.
📝 Abstract
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Problem

Research questions and friction points this paper is trying to address.

multi-teacher distillation
on-policy distillation
capability seesaw
cross-task interference
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Subspace Projection
Task Ordering
Conflict Probe
Geometry-Aware Optimization
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1
Y
Youzhi Liu
Ant Group
R
Ruobing Zheng
Ant Group
B
Boyuan Tong
Ant Group
T
Tianqi Li
Ant Group
P
Pingqi Li
Ant Group
H
Hanbo Bi
Ant Group
Yi Yuan
Yi Yuan
NetEase Fuxi AI Lab
deep learningcomputer vision
Jingdong Chen
Jingdong Chen
Senior Staff Algorithm Engineer, Ant Group
Computer VisionMultimodal