🤖 AI Summary
This study addresses the inefficiency of full adversarial fine-tuning and the poor cross-task transferability of robustness when adapting frozen speech foundation models to downstream tasks. To this end, we propose an efficient robust adaptation method built upon Wav2Vec2, HuBERT, and WavLM. By hierarchically stabilizing hidden-layer representations and fixing mixing weights, our approach decouples robustness from task-specific adaptation. It subsequently optimizes only the linear classifier boundary, enabling robust transfer without requiring downstream adversarial examples. Experimental results demonstrate that the proposed method improves robust accuracy by an average of 46.4 percentage points across four tasks, while incurring a marginal clean accuracy degradation of merely 1.1%.
📝 Abstract
Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial examples and updating the backbone for every task sacrifices that efficiency. We ask whether robustness can instead be learned before future tasks are known. For a frozen backbone and linear classifier, robustness can be understood through the interaction between representation stability and decision-boundary margin. This leads directly to our design: we stabilize representations across the hidden layers, rather than only the final layer, while preserving clean representations; after clean adaptation selects the layer mixture, we keep it fixed and enlarge only the classifier margin, without downstream adversarial examples. We evaluate Wav2Vec2, HuBERT, and WavLM Large on four tasks under adaptive 30 dB attacks. Across 12 backbone-task pairs, hierarchical robustification improves robust accuracy by 46.4 pp, while margin refinement adds 4.0 pp for 1.1 pp of clean accuracy. Code and configurations are available at https://github.com/arefmousavi/hierarchical-robust-sfm.