🤖 AI Summary
This study addresses the security vulnerabilities in fine-tuned automatic speech recognition (ASR) models arising from adversarial perturbation transferability via shared public base models. To enhance ASR robustness under black-box deployment, this work proposes TransferBreaker, a unified defense framework that effectively suppresses adversarial transferability. The method innovatively integrates ensemble-based adversarial fine-tuning, latent Jacobian regularization, and mixed gradient interpolation into a cohesive mechanism. Experimental evaluations demonstrate that TransferBreaker significantly reduces the adversarial word error rate on multilingual large models from 92.6% to 27.8%, thoroughly validating both its theoretical soundness and practical utility for securing deployed speech recognition systems against transferable adversarial attacks.
📝 Abstract
Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at https://github.com/rohban-lab/TransferBreaker.