anonymize speaker identity

Designs and implements algorithms and processing pipelines that remove or conceal speaker-identifying characteristics in speech audio—such as timbre and spectral-envelope cues—while retaining linguistic content, prosodic contours, and timing. Builds extraction modules to isolate target speakers from mixtures for anonymization, extends methods to multi‑speaker and child‑adapted scenarios, and analyzes privacy–utility tradeoffs and preservation of conversational structure in overlapping or two‑speaker conditions.

anonymizespeakeridentity

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

SecureSpeech: Prompt-based Speaker and Content Protection

Jul 10, 2025
BS
Belinda Soh Hui Hui
🏛️ Sinagpore Institute of Technology | Duke Kunshan University | National Institute of Informatics

To address dual privacy leakage—speaker identity and semantic content—in speech, this paper proposes the first end-to-end prompt-driven speech anonymization framework. Methodologically, it achieves speaker de-identification via controllable speaker descriptors, identifies and replaces sensitive semantic content using named entity recognition and large language models, and synthesizes high-fidelity anonymous speech via text-to-speech. Its key contributions are: (1) the first joint, controllable anonymization of both identity and semantics; (2) a descriptor modulation mechanism that explicitly models the privacy–utility trade-off; and (3) systematic analysis revealing how descriptor selection induces inherent biases between privacy protection and speech usability. Experiments demonstrate strong efficacy: speaker identification accuracy drops below 15%, while ASR accuracy remains above 92% and MOS for naturalness reaches ≈4.1—confirming both effectiveness and practicality.

Anonymizes sensitive content in speech using NLP modelsBalances privacy protection with speech quality and utilityProtects speaker identity from theft and re-identification

Target speaker anonymization in multi-speaker recordings

Oct 10, 2025
NT
Natalia Tomashenko
🏛️ Université de Lorraine | CNRS | Inria | Loria | National Institute of Informatics

This work addresses the underexplored problem of anonymizing only the target speaker—e.g., the customer in call-center dialogues—within multi-speaker conversational speech. We propose the first fine-grained target-speaker anonymization framework for multi-speaker audio. Our method integrates speaker separation, target-directed voiceprint feature perturbation, and context-aware speech reconstruction to precisely conceal the identity of the specified speaker. Innovatively, we design a joint privacy-utility evaluation metric to overcome the limitations of existing metrics in multi-speaker scenarios. Experiments demonstrate that our approach significantly enhances target-speaker identity protection—reducing privacy leakage by over 60%—while preserving speech intelligibility and conversational coherence. Degradations in speech quality (measured by PESQ) and ASR accuracy remain within acceptable bounds, confirming a favorable privacy–utility trade-off.

Addressing limitations of conventional single-speaker anonymization methodsAnonymizing target speakers in multi-speaker conversational recordingsDeveloping evaluation metrics for privacy and utility preservation

Enforcing Speech Content Privacy in Environmental Sound Recordings using Segment-wise Waveform Reversal

Jul 11, 2025
MT
Modan Tailleur
🏛️ Nantes Université | École Centrale Nantes | Université Gustave Eiffel | ENSA Nantes

Embedding intelligible speech in environmental sound recordings poses severe privacy risks, hindering data sharing and reuse. To address this, we propose a waveform-level privacy-preserving method based on segment-wise polarity inversion: speech segments are first precisely localized via voice activity detection and speaker separation; then, their local waveforms undergo polarity inversion, followed by randomized concatenation to enhance resistance against adversarial reconstruction. Crucially, this approach degrades speech intelligibility while preserving non-speech acoustic scene structure and overall audio fidelity. Experiments on a synthetic dataset show a 97.9% word error rate (WER), only a 2.7% drop in sound source classification accuracy, and a Fréchet Audio Distance of 1.40—substantially outperforming baseline methods. To our knowledge, this is the first work to combine waveform-level perturbation with randomized reassembly for environmental audio privacy protection, achieving a balanced trade-off among strong privacy guarantees, high perceptual quality, and downstream task utility.

Distort speech via segment-wise waveform reversalPreserve acoustic scene integrity and audio qualityProtect speech privacy in environmental recordings

A Benchmark for Multi-speaker Anonymization

Jul 08, 2024
XM
Xiaoxiao Miao
🏛️ Singapore Institute of Technology | National University of Singapore | National Institute of Informatics

Existing voice anonymization methods primarily target single-speaker scenarios and exhibit insufficient privacy protection in multi-speaker settings—especially under speech overlap. This work presents the first systematic study of multi-speaker voice anonymization, establishing the first dedicated benchmark encompassing task formalization, evaluation protocols, baseline systems, and privacy leakage analysis. We propose a dialogue-level dual-objective vector anonymization framework that simultaneously preserves inter-speaker relational structure and enhances individual speaker distinguishability. Technically, it integrates spectral-clustering-based speaker diarization, disentangled anonymization, selective anonymizers, and two dialogue-level speaker vector optimization strategies. Experiments on both non-overlapping simulated and real-world datasets demonstrate substantial reductions in privacy leakage risk, alongside improvements in speech intelligibility and speaker discriminability.

Addresses privacy leakage in overlapping conversationsDevelops benchmark for multi-speaker voice anonymizationEnhances utility with speaker vector anonymization methods

Adapting General Disentanglement-Based Speaker Anonymization for Enhanced Emotion Preservation

Aug 12, 2024
XM
Xiaoxiao Miao
🏛️ Singapore Institute of Technology | Chinese Academy of Sciences | National Institute of Informatics | Inria

This study addresses severe emotional distortion in disentangled speaker anonymization. We propose a novel method that jointly ensures strong privacy protection and high emotional fidelity. Methodologically, we (1) design a pre-trained emotion encoder to explicitly disentangle emotion representations from speaker identity; (2) introduce the first SVM-based emotion boundary modeling and directional embedding-space compensation mechanism, enabling fine-grained adjustment of anonymized speaker embeddings along emotion gradients; and (3) integrate dual-path compensation—encoder-level feature fusion and post-hoc refinement. Experiments demonstrate that our approach achieves state-of-the-art speaker unidentifiability (ASR-based ID error >95%) while improving emotion recognition accuracy by 12.7% over existing disentanglement-based anonymization methods. Moreover, the framework is extensible to controllable preservation of other paralinguistic attributes, such as age and accent.

Adapt speaker anonymization to preserve emotional cuesBalance privacy protection and emotion retentionExtend system to preserve other paralinguistic attributes

Latest Papers

What's happening recently
View more

This work addresses the significant performance degradation of existing voice anonymization systems—primarily designed for adult speech—when applied to children’s voices, which struggle to balance privacy preservation with speech utility. To bridge this gap, the study proposes the first systematic self-supervised learning (SSL)-based anonymization framework tailored specifically for child speech. By performing domain adaptation on the MyST child corpus and extending the approach to both single-speaker and two-speaker mixture scenarios, the method integrates target speaker extraction with explicit modeling of child-specific vocal characteristics. This enables effective identity privacy protection while substantially improving speech intelligibility, perceptual quality, and conversational naturalness, thereby achieving child-centric, high-fidelity voice anonymization.

child speechmulti-speakerspeaker privacy

This work addresses the challenge of voice anonymization by proposing a method that effectively removes speaker identity while preserving linguistic content fidelity, without relying on complex waveform reconstruction losses or explicit speaker embeddings. The approach leverages a frozen wav2vec 2.0 encoder to extract content embeddings, which are then vector-quantized and fed into a HiFi-GAN vocoder to synthesize high-quality speech. An adversarial speaker classification branch with a gradient reversal layer is introduced to deliberately confuse identity information. Requiring only content embedding alignment and adversarial training—without waveform-level losses or speaker embedding mappings—the system achieves strong performance: a word error rate of 2.53% and an anonymization equal error rate (EER) of 13.39% on the Voice Privacy Challenge (VPC) benchmark, placing it among the top-tier systems, while also unexpectedly retaining emotional characteristics with a UAR of 43.91%, yielding clear and natural-sounding output.

content preservationembedding matchingspeaker identity removal

This work addresses the lack of systematic evaluation of privacy leakage at the speaker attribute level in existing voice anonymization methods. It proposes the first privacy assessment framework grounded in speaker attributes, which systematically quantifies privacy risks by comparing ground-truth attributes, attributes inferred from original speech, and those inferred from anonymized speech. The framework integrates speaker uniqueness metrics with single-utterance attack error rate analysis to evaluate residual identifiability. The study reveals that even in the presence of attribute inference errors, inferred attributes can still pose significant privacy threats, exposing a critical gap in current anonymization techniques’ ability to protect attribute-level information. These findings underscore the necessity for future research to incorporate attribute-aware threat modeling and develop corresponding defense mechanisms.

attribute inferenceprivacy riskspeaker anonymity

This study addresses the limitations of existing deepfake speech detection methods, which struggle to model speaker-specific pronunciation patterns and thus inadequately protect high-profile individuals. To overcome this, the authors propose a phoneme-based voice profiling (PVP) framework that pioneers a shift from utterance-level to phoneme-level fine-grained modeling. By employing a lightweight Gaussian Mixture Model (GMM), PVP captures speaker-specific acoustic distributions of individual phonemes, enabling the construction of an interpretable, personalized voiceprint using only a small amount of genuine speech—without requiring any synthetic or spoofed samples for training. The work also introduces the first Chinese deepfake dataset featuring public figures, on which the proposed method significantly outperforms state-of-the-art general-purpose detectors, substantially reducing the equal error rate (EER) and supporting phoneme-level forensic analysis.

deepfake detectionphonemespeaker-specific

Traditional speaker anonymization evaluation relies on aggregate metrics, overlooking substantial inter-speaker variations in re-identification risk. This study conducts a large-scale, speaker-level privacy analysis of nearly 5,000 speakers under worst-case scenarios, systematically assessing re-identification risks across diverse combinations of anonymization methods, adversary architectures, and utterance lengths using linkability-based metrics. The findings reveal that re-identification difficulty is not determined by inherent speaker characteristics but emerges from the interaction among the anonymization scheme, adversary capability, and available speech duration. Crucially, the composition of easily or hardly identifiable speaker groups shifts markedly with system configuration, challenging the assumption of static privacy risk and underscoring the necessity of conditioning privacy evaluations on specific attack models and anonymization settings.

linkabilityper-speaker analysisprivacy evaluation

Hot Scholars

XM

Xiaoxiao Miao

Duke Kunshan University
Speech PrivacySpeaker and Language IdentificationSpeech Synthesis
FT

Francisco Teixeira

INESC-ID / IST, University of Lisbon
Speech ProcessingPrivacyMachine Learning
EH

Enno Hermann

Postdoc, IDIAP Research Institute, Switzerland
Speech RecognitionSpeech SynthesisNatural Language ProcessingMachine Learning
TB

Tom Bäckström

Aalto University
privacy and security in speech communicationspeech enhancementacoustic sensor networksspeech and audio coding