Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion

πŸ“… 2026-07-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the growing threat of voice conversion (VC) technologies being exploited for biometric spoofing, which undermines speaker identity verification in forensic applications. To recover the original speaker identity from VC-manipulated speech, the paper proposes TRIDENT, a novel tripartite collaborative framework comprising a shared backbone extractor and two auxiliary branches dedicated respectively to identifying the VC method type and extracting the target speaker’s characteristics. By incorporating prior assumptions about mainstream VC model families and leveraging multi-task learning with disentangled latent representations, TRIDENT effectively decouples confounding factors introduced by conversion. The approach achieves a 90.99% source speaker identification accuracy across seven state-of-the-art VC methods and demonstrates robustness under challenging conditions, including telephone-bandwidth channels, unseen languages, and adaptive scenarios.
πŸ“ Abstract
Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.
Problem

Research questions and friction points this paper is trying to address.

voice conversion
speaker identity recovery
biometric security
forensic audio analysis
source speaker identification
Innovation

Methods, ideas, or system contributions that make the work stand out.

voice conversion
speaker identity recovery
forensic audio analysis
multi-branch architecture
biometric security
πŸ”Ž Similar Papers
No similar papers found.