🤖 AI Summary
This study addresses the inherent conflict between pretrained speaker embeddings and streaming acoustic generation constraints in real-time voice conversion by proposing the VOSSA framework. Rather than employing an independent speaker encoder, this method extracts and aggregates speaker information from the intermediate layers of a content encoder. It further incorporates attention-based statistical pooling to optimize acoustic dynamics in streaming scenarios and achieves end-to-end optimization through multi-task joint training. Experimental results demonstrate that VOSSA significantly enhances fundamental frequency dynamics and vowel cues. While maintaining low error rates, the proposed framework effectively improves the naturalness, speaker similarity, and intelligibility of the synthesized speech.
📝 Abstract
Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.