🤖 AI Summary
Understanding how self-supervised speech representation models encode speaker-specific paralinguistic attributes—such as pitch, speaking rate, and energy—remains underexplored, especially in comparison to dedicated speaker embedding models.
Method: We systematically analyze intermediate-layer representations from HuBERT, WavLM, and Wav2Vec 2.0, alongside CAM++, using probe-based linear classification on quantitatively annotated acoustic features.
Contribution/Results: Our study is the first to comparatively evaluate multi-dimensional interpretability across layers and models. We find that deeper SSL layers achieve superior joint discriminability across paralinguistic dimensions; CAM++ excels specifically in energy classification; and intermediate layers naturally integrate acoustic and paralinguistic information, facilitating feature disentanglement. These findings establish a hierarchical, interpretable representation foundation for speaker verification and text-to-speech synthesis.
📝 Abstract
This study explores speaker-specific features encoded in speaker embeddings and intermediate layers of speech self-supervised learning (SSL) models. By utilising a probing method, we analyse features such as pitch, tempo, and energy across prominent speaker embedding models and speech SSL models, including HuBERT, WavLM, and Wav2vec 2.0. The results reveal that speaker embeddings like CAM++ excel in energy classification, while speech SSL models demonstrate superior performance across multiple features due to their hierarchical feature encoding. Intermediate layers effectively capture a mix of acoustic and para-linguistic information, with deeper layers refining these representations. This investigation provides insights into model design and highlights the potential of these representations for downstream applications, such as speaker verification and text-to-speech synthesis, while laying the groundwork for exploring additional features and advanced probing methods.