🤖 AI Summary
This work addresses the challenge of voice anonymization by proposing a method that effectively removes speaker identity while preserving linguistic content fidelity, without relying on complex waveform reconstruction losses or explicit speaker embeddings. The approach leverages a frozen wav2vec 2.0 encoder to extract content embeddings, which are then vector-quantized and fed into a HiFi-GAN vocoder to synthesize high-quality speech. An adversarial speaker classification branch with a gradient reversal layer is introduced to deliberately confuse identity information. Requiring only content embedding alignment and adversarial training—without waveform-level losses or speaker embedding mappings—the system achieves strong performance: a word error rate of 2.53% and an anonymization equal error rate (EER) of 13.39% on the Voice Privacy Challenge (VPC) benchmark, placing it among the top-tier systems, while also unexpectedly retaining emotional characteristics with a UAR of 43.91%, yielding clear and natural-sounding output.
📝 Abstract
The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization and a HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddings of the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.