🤖 AI Summary
This study addresses the degradation in model performance caused by fixed distillation losses when merging speech and music encoders. To overcome this limitation, it proposes a teacher-student knowledge distillation framework based on correlation loss, integrating weight interpolation and channel permutation techniques to optimize student model representations. The research reveals the critical influence of distillation losses on model mergeability, demonstrating that performance improvements stem from enhanced representation quality rather than specific merging mechanisms. Empirically, the proposed method achieves superior endpoint performance across most comparative experiments, significantly reduces the reliance on channel rearrangement, and effectively improves performance on speech-related tasks.
📝 Abstract
A single compact encoder for both speech and music removes the need to maintain a separate model per domain. A practical recipe distils a speech teacher and a music teacher into two small students, then merges them into one model. Two merging algorithms exist: interpolating the student weights from a shared initialisation, or permuting their channels into alignment before averaging. Prior work fixed the distillation loss in both, leaving open whether the loss itself affects how well the students merge. We show that it does. Holding architecture and evaluation fixed, we replace the DistilHuBERT loss \Lkd{} with a correlation-based loss \Lcl{}. Under weight interpolation, \Lcl{} students stay closer to their own better endpoint in $17$ of $18$ comparisons and lead on the speech tasks at every interpolated weight. Under activation permutation, they need fewer channels rearranged at three layers, their matched channels correlate more strongly, and they again lead on speech. The two merging algorithms share no mechanism, yet both improve under the same loss substitution. This suggests that the gain lies in the representation \Lcl{} produces rather than in either algorithm.