Quantifying Source Speaker Leakage in One-to-One Voice Conversion

📅 2024-09-25
🏛️ Biometrics and Electronic Signatures
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the privacy risk of source speaker identity leakage in one-to-one voice conversion, providing the first systematic quantification of its identifiability. Under a white-box assumption, we construct a multi-accent parallel corpus and generate high-fidelity converted speech using the HiFi-GAN vocoder. We then propose a classification-confidence-based source identity inference framework to interpretably model residual source characteristics and assess leakage severity. Our method significantly narrows the candidate set of potential source speakers—achieving an average reduction rate of 62%—and establishes the first actionable, reproducible metric for identity privacy compliance in voice synthesis services. The core contributions are: (1) establishing a quantitative paradigm for privacy leakage in voice conversion; (2) explicitly characterizing the boundary of source identity information preserved in converted speech; and (3) enabling regulatory implementation and ethical accountability by grounding privacy assessment in empirically measurable metrics.

Technology Category

Machine Learning: PrivacyNatural Language Processing: Ethics — Bias, Fairness, Transparency & PrivacyComputer Vision: Bias, Fairness & Privacy

Application Category

Security and Privacy: Large-scale security measurementsUser Modeling, Personalization and Recommendation: User privacy protection in personalized systemsResponsible Web: Data and user privacy-enhancing technologies for the Web
📝 Abstract
Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker’s identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a ‘worst-case’ white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers’ data.
Problem

Research questions and friction points this paper is trying to address.

Quantify source speaker leakage in voice conversion
Assess confidence in identifying original speakers
Ensure privacy in synthetic voice data usage
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses HiFi-GAN vocoder for voice conversion
Quantifies source speaker identity leakage
Analyzes multi-accented parallel utterance corpus
🔎 Similar Papers
No similar papers found.