🤖 AI Summary
This study addresses the lack of a unified quantitative evaluation framework for assessing the disentanglement of speaker identity and prosody in speech content representations. To this end, we construct a generative model relying solely on a single representation to systematically compare self-supervised learning (SSL) features and supervised tokens across content, identity, and prosody dimensions. Our findings reveal that disentanglement efficacy is governed by the interplay between training objectives and information capacity, rather than being determined exclusively by supervisory signals. Furthermore, we identify two distinct representational paradigms: high-fidelity reconstruction and strong disentanglement. We demonstrate that, under constrained capacity, supervised representations can effectively isolate speaker identity. These insights provide a novel theoretical foundation for advancing speech representation learning.
📝 Abstract
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.