Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of general capabilities in multi-teacher online distillation, where student models drift from their initialization. To mitigate this issue, we propose a slow-fast dual-model coupling architecture. Specifically, an exponential moving average (EMA) is employed to construct a slow model that serves as a capability anchor. Furthermore, we introduce a novel orthogonal projection mechanism in the log-probability space, which precisely filters out conflicting gradient updates that push the student away from the slow model while preserving both alignment and orthogonal components. This framework effectively balances domain-specific learning with the retention of general capabilities, significantly alleviating knowledge interference. Empirical results demonstrate that our approach substantially enhances specialized multimodal performance while effectively suppressing degradation on general benchmarks.
📝 Abstract
Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
Problem

Research questions and friction points this paper is trying to address.

Multi-teacher distillation
Capability interference
Multimodal large language models
Capability preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Teacher On-Policy Distillation
Slow-Fast Architecture
Exponential Moving Average
Log-Probability Space Projection
Capability Preservation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
X
Xiaofei Yin
Ant Security Lab, Ant Group, Shanghai, China
T
Tong Chu
Ant Security Lab, Ant Group, Shanghai, China
Jiyuan Fu
Jiyuan Fu
Fudan University
Jun Lan
Jun Lan
Ant Group
Shuheng Zhou
Shuheng Zhou
Professor, University of California, Riverside
Machine learningHigh dimensional statisticsDifferential PrivacyAlgorithms
H
Huijia Zhu
Ant Security Lab, Ant Group, Shanghai, China