🤖 AI Summary
This study addresses the degradation of general capabilities in multi-teacher online distillation, where student models drift from their initialization. To mitigate this issue, we propose a slow-fast dual-model coupling architecture. Specifically, an exponential moving average (EMA) is employed to construct a slow model that serves as a capability anchor. Furthermore, we introduce a novel orthogonal projection mechanism in the log-probability space, which precisely filters out conflicting gradient updates that push the student away from the slow model while preserving both alignment and orthogonal components. This framework effectively balances domain-specific learning with the retention of general capabilities, significantly alleviating knowledge interference. Empirical results demonstrate that our approach substantially enhances specialized multimodal performance while effectively suppressing degradation on general benchmarks.
📝 Abstract
Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.