Score
Design and train representation-learning models that use adversarial objectives (for example, a speaker-classifier adversary) to remove or minimize speaker identity information from speech/audio embeddings; build the adversarial discriminators and encoders that enforce speaker-invariance. Evaluate and analyze the resulting embeddings for residual speaker-specific acoustic cues and for robustness to speaker-induced variability in downstream tasks (e.g., emotion or content classification).
To address the poor cross-model transferability of adversarial examples in automatic speech recognition (ASR) systems, this paper proposes an acoustic representation optimization method: for the first time, adversarial perturbations are constrained within a model-agnostic, low-level robust acoustic feature space, thereby unifying perturbation alignment and transferability. The method is plug-and-play, compatible with mainstream audio adversarial frameworks, and requires no modification to target models. Black-box attack experiments across three state-of-the-art ASR models demonstrate an average 32.7% improvement in transfer success rate, while strictly preserving perceptual fidelity of the original speech. Key contributions include: (1) establishing an acoustic-representation-driven paradigm for enhancing adversarial transferability; (2) achieving synergistic optimization of high transferability and high fidelity; and (3) providing a general, lightweight, and model-agnostic adversarial enhancement solution that requires no access to target model internals.
This study investigates the implicit encoding of speaker-intrinsic attributes—such as gender, age, fundamental frequency (F0), speaking rate, and utterance duration—in spoof embeddings for voice anti-spoofing detection, and examines how such encoding affects system robustness. Using probing classification, lightweight neural classifiers are trained on the ASVspoof 2019 Logical Access dataset to quantitatively assess the preservation of multidimensional speaker metadata and acoustic features within spoof embeddings. Results reveal, for the first time, that despite being optimized solely for spoof detection, spoof embeddings retain statistically significant information about gender, speaking rate, F0, and utterance duration—attributes known to influence robustness. This indicates an inherent speaker representation capability in spoof embeddings. The finding advances interpretability and generalization analysis of deep forgery-detection models and suggests that speaker-related information may serve as a critical underpinning for spoof detection robustness.
This work addresses the limited generalization of existing voice spoofing detection models in cross-domain scenarios, which often stems from their over-reliance on speaker identity cues at the expense of genuine spoofing artifacts. To mitigate this, the authors propose a speaker-label-free teacher–student framework that leverages a pre-trained speaker recognition model as the teacher. A gradient reversal layer steers the student network to learn speaker-invariant representations, while a variational information bottleneck is introduced to balance the suppression of speaker identity information against the preservation of spoofing-related cues. This approach achieves, for the first time, unsupervised learning of speaker-invariant representations for spoofing detection, effectively disentangling speaker and spoofing characteristics. Experiments across nine datasets demonstrate a relative 25.7% reduction in equal error rate (EER) compared to the MHFA baseline.
Phonetic-level perturbations in speech adversarial attacks—such as vowel centralization and consonant substitution—induce significant identity drift, jointly degrading both automatic speech recognition (ASR) and speaker verification (SV) systems. Method: We systematically analyze, from a phoneme-centric perspective, the root causes of speaker identity distortion in adversarial examples, proposing a novel phoneme-aware defense paradigm. Targeting DeepSpeech, we generate targeted adversarial samples and quantitatively evaluate their impact on transcription accuracy and speaker embedding distributions using both genuine and impostor speech. Experiments span 16 phonetically diverse target phrases. Contribution/Results: All phrases exhibit high transcription error rates and substantial speaker embedding shifts, confirming that phoneme-level perturbations constitute a synergistic threat to ASR and SV. This work establishes a theoretically grounded, interpretable framework for enhancing speech robustness against multi-task adversarial attacks.
To address speaker representation degradation and verification performance decline caused by emotional variability, this paper proposes an emotion-robust speaker representation learning framework. Methodologically, it introduces three novel components: (1) Copy-Paste speech data augmentation, (2) a cosine similarity–constrained loss function, and (3) an energy-based masking (EM) mechanism for emotion-aware spectral suppression. The EM module dynamically attenuates emotion-discriminative frequency bands using frame-level speech energy, thereby explicitly disentangling speaker identity from emotion-related variations and enhancing representation invariance. Evaluated on standard benchmarks, the proposed approach achieves a 19.29% relative reduction in equal error rate (EER) over baseline systems. Ablation studies confirm the individual efficacy and synergistic benefits of all components. This work provides a principled, reproducible methodology for building highly robust speaker verification systems under emotional variability.
This work addresses the limited transferability of existing black-box audio adversarial attacks and their vulnerability to waveform-level defenses. The authors propose a novel attack method that operates in the self-supervised learning (SSL) feature space: adversarial perturbations are crafted in the SSL-based acoustic-phonetic representation and then reconstructed into speech-like waveforms via a vocoder, thereby evading waveform-level defenses and enhancing cross-model transferability. Using only Whisper-small as a surrogate model, the proposed approach achieves a +26.6% absolute increase in Word Error Rate (WER) across multiple black-box automatic speech recognition (ASR) systems and demonstrates robustness against various defense mechanisms, yielding a +36.2% WER improvement—significantly outperforming current state-of-the-art methods.
This work addresses the limited generalization of current anti-spoofing systems when confronted with speech generated by unseen synthesis methods. To enhance robustness, the study introduces a Mixture-of-Experts (MoE) architecture into self-supervised speech representation models for the first time. Specifically, the standard feed-forward modules in key encoding layers are replaced with multiple expert networks, and a learnable per-layer gating mechanism is designed to enable collaborative modeling of complementary acoustic features while preserving the pre-trained representational capabilities. The proposed approach significantly improves generalization to previously unseen spoofed speech, reducing the macro-averaged equal error rate (EER) from 5.46% to 4.81% across 14 spoofing datasets—a relative improvement of 11.9%.
This work addresses the vulnerability of automatic speech recognition (ASR) systems to adversarial perturbations—distortions imperceptible to humans yet capable of inducing transcription errors. The authors propose a neural audio codec based on residual vector quantization (RVQ) that introduces a discrete bottleneck in the signal pathway to suppress adversarial noise while preserving linguistic content. Their analysis reveals a non-monotonic trade-off between quantization depth and robustness, demonstrating that intermediate RVQ depths optimally balance content fidelity and adversarial resilience. Notably, the study establishes, for the first time, a strong correlation between discrete codebook alterations and transcription errors. Experimental results show that the proposed method significantly reduces word error rates across multiple attack types, outperforming conventional compression-based defenses and maintaining robustness even under adaptive attacks.
This study addresses the unclear impact of speaker identity on embedding representations in existing voice anti-spoofing systems. To this end, it presents the first systematic evaluation and effective disentanglement of speaker-related factors from anti-spoofing performance. The work proposes two contrasting speaker-invariant modeling strategies: a joint multi-task learning approach and an explicit removal of speaker information via a gradient reversal layer. Evaluated across four benchmark datasets, the proposed methods demonstrate substantial improvements, reducing the average equal error rate by 17% and achieving up to a 48% reduction for the most challenging attack types (e.g., A11). These results significantly enhance the model’s generalization capability and robustness against diverse spoofing attacks.
This work addresses the performance degradation in multi-corpus joint training for anti-spoofing tasks, which often arises from dataset-specific biases leading to negative transfer. To mitigate this issue, the study introduces domain-invariant learning into self-supervised speech anti-spoofing models for the first time, proposing an Invariant Domain Feature Extraction (IDFE) framework. By integrating multi-task learning with gradient reversal layers, IDFE effectively disentangles corpus-specific information and enhances cross-dataset generalization. Experimental results across four mainstream anti-spoofing datasets demonstrate that the proposed method achieves a 20% relative reduction in average equal error rate compared to baseline models, substantially alleviating the instability commonly observed in multi-corpus training scenarios.