What Actually Makes Correlation-Based SSL Distillation Noise-Robust? A Mechanistic Correction

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study elucidates the true mechanisms underlying noise robustness in knowledge distillation for self-supervised speech models, correcting prevalent misconceptions regarding the role of auto-correlation terms. Through feature-space mechanistic analysis employing unbiased full-dimensional probing, Pearson variance decomposition, and causally controlled experiments, this work isolates and verifies the individual contributions of each component for the first time. We demonstrate that the diagonal of the cross-correlation matrix is pivotal for enhancing noise robustness, whereas auto-correlation terms primarily optimize downstream task accuracy. Accordingly, we propose a design principle enabling independent weight modulation of these two components and establish best practices for independently sampling teacher and student noise. This approach reduces noise classification accuracy from 76.98% to 55.02%, offering novel theoretical guidance for robustness-oriented distillation.
📝 Abstract
Self-supervised learning (SSL) speech models are accurate but large. Knowledge distillation compresses them, but the student loses the teacher's noise robustness. Correlation-based distillation addresses this with two terms: cross-correlation aligning student and teacher's representations, and self-correlation decorrelating the student's features. Both the original method and De'HuBERT credited the self-correlation term without isolating it. We show the opposite. On LibriSpeech-100 distillation with held-out CHiME-3 noise at 10\,dB, an unbiased full-dimensional probe, a Pearson-variance decomposition, a same-noise causal control, and a per-dimension analysis identify the cross-correlation diagonal as the mechanism that encourages noise invariance, lowering noise-classification accuracy from 76.98\% to 55.02\%, whereas adding the self-correlation term back leaves it at 55.55\%. The self-correlation term instead reorganises the feature space and improves accuracy on the clean downstream tasks, but removes essentially no noise. Across nine speech and music tasks this yields a concrete design rule for distillation: weight the cross-correlation term for noise robustness, tune the self-correlation term for better downstream performance, and sample teacher and student noise independently.
Problem

Research questions and friction points this paper is trying to address.

Self-supervised learning
Knowledge distillation
Noise robustness
Correlation-based distillation
Speech models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Knowledge Distillation
Self-Supervised Learning
Noise Robustness
Cross-Correlation
Mechanistic Interpretability
🔎 Similar Papers
No similar papers found.