speaker-adversarial representation learning

Design and train representation-learning models that use adversarial objectives (for example, a speaker-classifier adversary) to remove or minimize speaker identity information from speech/audio embeddings; build the adversarial discriminators and encoders that enforce speaker-invariance. Evaluate and analyze the resulting embeddings for residual speaker-specific acoustic cues and for robustness to speaker-induced variability in downstream tasks (e.g., emotion or content classification).

speaker-adversarialrepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.68
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Boosting the Transferability of Audio Adversarial Examples with Acoustic Representation Optimization

Mar 25, 2025
WJ
Weifei Jin
🏛️ Beijing University of Posts and Telecommunications

To address the poor cross-model transferability of adversarial examples in automatic speech recognition (ASR) systems, this paper proposes an acoustic representation optimization method: for the first time, adversarial perturbations are constrained within a model-agnostic, low-level robust acoustic feature space, thereby unifying perturbation alignment and transferability. The method is plug-and-play, compatible with mainstream audio adversarial frameworks, and requires no modification to target models. Black-box attack experiments across three state-of-the-art ASR models demonstrate an average 32.7% improvement in transfer success rate, while strictly preserving perceptual fidelity of the original speech. Key contributions include: (1) establishing an acoustic-representation-driven paradigm for enhancing adversarial transferability; (2) achieving synergistic optimization of high transferability and high fidelity; and (3) providing a general, lightweight, and model-agnostic adversarial enhancement solution that requires no access to target model internals.

Addressing lack of model-specific information in real-world attack scenariosEnhancing transferability of audio adversarial examples across ASR modelsOptimizing perturbations using low-level acoustic representations for consistency

Explaining Speaker and Spoof Embeddings via Probing

Dec 24, 2024
XL
Xuechen Liu
🏛️ National Institute of Informatics | Institute for Advancing Intelligence | University of Eastern Finland

This study investigates the implicit encoding of speaker-intrinsic attributes—such as gender, age, fundamental frequency (F0), speaking rate, and utterance duration—in spoof embeddings for voice anti-spoofing detection, and examines how such encoding affects system robustness. Using probing classification, lightweight neural classifiers are trained on the ASVspoof 2019 Logical Access dataset to quantitatively assess the preservation of multidimensional speaker metadata and acoustic features within spoof embeddings. Results reveal, for the first time, that despite being optimized solely for spoof detection, spoof embeddings retain statistically significant information about gender, speaking rate, F0, and utterance duration—attributes known to influence robustness. This indicates an inherent speaker representation capability in spoof embeddings. The finding advances interpretability and generalization analysis of deep forgery-detection models and suggests that speaker-related information may serve as a critical underpinning for spoof detection robustness.

Acoustic Feature AnalysisSpeaker VerificationVoice Fraud Detection

This work addresses the limited generalization of existing voice spoofing detection models in cross-domain scenarios, which often stems from their over-reliance on speaker identity cues at the expense of genuine spoofing artifacts. To mitigate this, the authors propose a speaker-label-free teacher–student framework that leverages a pre-trained speaker recognition model as the teacher. A gradient reversal layer steers the student network to learn speaker-invariant representations, while a variational information bottleneck is introduced to balance the suppression of speaker identity information against the preservation of spoofing-related cues. This approach achieves, for the first time, unsupervised learning of speaker-invariant representations for spoofing detection, effectively disentangling speaker and spoofing characteristics. Experiments across nine datasets demonstrate a relative 25.7% reduction in equal error rate (EER) compared to the MHFA baseline.

out-of-domain generalizationspeaker biasspeaker-invariant representation

Impact of Phonetics on Speaker Identity in Adversarial Voice Attack

Sep 18, 2025
DK
Daniyal Kabir Dar
🏛️ Michigan State University

Phonetic-level perturbations in speech adversarial attacks—such as vowel centralization and consonant substitution—induce significant identity drift, jointly degrading both automatic speech recognition (ASR) and speaker verification (SV) systems. Method: We systematically analyze, from a phoneme-centric perspective, the root causes of speaker identity distortion in adversarial examples, proposing a novel phoneme-aware defense paradigm. Targeting DeepSpeech, we generate targeted adversarial samples and quantitatively evaluate their impact on transcription accuracy and speaker embedding distributions using both genuine and impostor speech. Experiments span 16 phonetically diverse target phrases. Contribution/Results: All phrases exhibit high transcription error rates and substantial speaker embedding shifts, confirming that phoneme-level perturbations constitute a synergistic threat to ASR and SV. This work establishes a theoretically grounded, interpretable framework for enhancing speech robustness against multi-task adversarial attacks.

Analyzing phonetic basis of adversarial audio perturbationsExploring phonetic distortions causing transcription and identity errorsInvestigating impact of perturbations on speaker identity verification

Learning Emotion-Invariant Speaker Representations for Speaker Verification

Apr 14, 2024
JT
Jingguang Tian
🏛️ Hithink RoyalFlush AI Research Institute

To address speaker representation degradation and verification performance decline caused by emotional variability, this paper proposes an emotion-robust speaker representation learning framework. Methodologically, it introduces three novel components: (1) Copy-Paste speech data augmentation, (2) a cosine similarity–constrained loss function, and (3) an energy-based masking (EM) mechanism for emotion-aware spectral suppression. The EM module dynamically attenuates emotion-discriminative frequency bands using frame-level speech energy, thereby explicitly disentangling speaker identity from emotion-related variations and enhancing representation invariance. Evaluated on standard benchmarks, the proposed approach achieves a 19.29% relative reduction in equal error rate (EER) over baseline systems. Ablation studies confirm the individual efficacy and synergistic benefits of all components. This work provides a principled, reproducible methodology for building highly robust speaker verification systems under emotional variability.

Enhancing speaker verification robustness against emotional variabilityImproving emotion-invariant features via data augmentation and maskingReducing emotion correlation in speaker representations

Latest Papers

What's happening recently
View more

This work addresses the limited transferability of existing black-box audio adversarial attacks and their vulnerability to waveform-level defenses. The authors propose a novel attack method that operates in the self-supervised learning (SSL) feature space: adversarial perturbations are crafted in the SSL-based acoustic-phonetic representation and then reconstructed into speech-like waveforms via a vocoder, thereby evading waveform-level defenses and enhancing cross-model transferability. Using only Whisper-small as a surrogate model, the proposed approach achieves a +26.6% absolute increase in Word Error Rate (WER) across multiple black-box automatic speech recognition (ASR) systems and demonstrates robustness against various defense mechanisms, yielding a +36.2% WER improvement—significantly outperforming current state-of-the-art methods.

adversarial attacksautomatic speech recognitionblack-box transferability

This work addresses the limited generalization of current anti-spoofing systems when confronted with speech generated by unseen synthesis methods. To enhance robustness, the study introduces a Mixture-of-Experts (MoE) architecture into self-supervised speech representation models for the first time. Specifically, the standard feed-forward modules in key encoding layers are replaced with multiple expert networks, and a learnable per-layer gating mechanism is designed to enable collaborative modeling of complementary acoustic features while preserving the pre-trained representational capabilities. The proposed approach significantly improves generalization to previously unseen spoofed speech, reducing the macro-averaged equal error rate (EER) from 5.46% to 4.81% across 14 spoofing datasets—a relative improvement of 11.9%.

anti-spoofingrobustnessspeech synthesis

This work addresses the vulnerability of automatic speech recognition (ASR) systems to adversarial perturbations—distortions imperceptible to humans yet capable of inducing transcription errors. The authors propose a neural audio codec based on residual vector quantization (RVQ) that introduces a discrete bottleneck in the signal pathway to suppress adversarial noise while preserving linguistic content. Their analysis reveals a non-monotonic trade-off between quantization depth and robustness, demonstrating that intermediate RVQ depths optimally balance content fidelity and adversarial resilience. Notably, the study establishes, for the first time, a strong correlation between discrete codebook alterations and transcription errors. Experimental results show that the proposed method significantly reduces word error rates across multiple attack types, outperforming conventional compression-based defenses and maintaining robustness even under adaptive attacks.

adversarial robustnesscapacity-robustness trade-offdiscrete bottleneck

This study addresses the unclear impact of speaker identity on embedding representations in existing voice anti-spoofing systems. To this end, it presents the first systematic evaluation and effective disentanglement of speaker-related factors from anti-spoofing performance. The work proposes two contrasting speaker-invariant modeling strategies: a joint multi-task learning approach and an explicit removal of speaker information via a gradient reversal layer. Evaluated across four benchmark datasets, the proposed methods demonstrate substantial improvements, reducing the average equal error rate by 17% and achieving up to a 48% reduction for the most challenging attack types (e.g., A11). These results significantly enhance the model’s generalization capability and robustness against diverse spoofing attacks.

embeddingsmulti-task learningspeaker identity

This work addresses the performance degradation in multi-corpus joint training for anti-spoofing tasks, which often arises from dataset-specific biases leading to negative transfer. To mitigate this issue, the study introduces domain-invariant learning into self-supervised speech anti-spoofing models for the first time, proposing an Invariant Domain Feature Extraction (IDFE) framework. By integrating multi-task learning with gradient reversal layers, IDFE effectively disentangles corpus-specific information and enhances cross-dataset generalization. Experimental results across four mainstream anti-spoofing datasets demonstrate that the proposed method achieves a 20% relative reduction in average equal error rate compared to baseline models, substantially alleviating the instability commonly observed in multi-corpus training scenarios.

domain biasgeneralizationmulti-corpus training

Hot Scholars

ZJ

Zeyu Jin

Adobe Research
Speech and audio processingDeep Learning
KL

Kyogu Lee

Professor, Seoul National University
Audio Signal ProcessingMachine LearningComputer AuditionAuditory/Music Perception & Cognition
SH

Shujie Hu

The Chinese University of Hong Kong
Speech ProcessingMLLM
HC

Hongyang Chen

SUN YAT-SEN UNIVERSITY
SDNCloud ComputingMicroserviceAIOps