🤖 AI Summary
This study addresses the privacy risk of source speaker identity leakage in one-to-one voice conversion, providing the first systematic quantification of its identifiability. Under a white-box assumption, we construct a multi-accent parallel corpus and generate high-fidelity converted speech using the HiFi-GAN vocoder. We then propose a classification-confidence-based source identity inference framework to interpretably model residual source characteristics and assess leakage severity. Our method significantly narrows the candidate set of potential source speakers—achieving an average reduction rate of 62%—and establishes the first actionable, reproducible metric for identity privacy compliance in voice synthesis services. The core contributions are: (1) establishing a quantitative paradigm for privacy leakage in voice conversion; (2) explicitly characterizing the boundary of source identity information preserved in converted speech; and (3) enabling regulatory implementation and ethical accountability by grounding privacy assessment in empirically measurable metrics.
📝 Abstract
Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker’s identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a ‘worst-case’ white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers’ data.