🤖 AI Summary
This study investigates the interpretability of speaker embedding spaces by examining their mapping relationships with conventional acoustic attributes—including phonetic features, gender, and age. Using a 10,000-speaker corpus, we systematically model correlations between embedding vectors from three state-of-the-art systems (x-vector, ECAPA-TDNN, and ResNet34) and nine interpretable acoustic parameters, employing principal component analysis and linear regression. Results show that these nine parameters collectively explain over 50% of the embedding variance—matching the explanatory power of the top-10 principal components. The embedding space robustly encodes gender information, exhibiting strong implicit gender discriminability; however, it demonstrates limited capacity to represent age-related variation. To our knowledge, this is the first work to quantitatively characterize the acoustic interpretability boundary of speaker embeddings, providing empirical grounding for interpretable speech representation learning and bias analysis, as well as concrete directions for architectural and training-level refinement.
📝 Abstract
Speaker embeddings are widely used in speaker verification systems and other applications where it is useful to characterise the voice of a speaker with a fixed-length vector. These embeddings tend to be treated as "black box" encodings, and how they relate to conventional acoustic and phonetic dimensions of voices has not been widely studied. In this paper we investigate how state-of-the-art speaker embedding systems represent the acoustic characteristics of speakers as described by conventional acoustic descriptors, age, and gender. Using a large corpus of 10,000 speakers and three embedding systems we show that a small set of 9 acoustic parameters chosen to be "interpretable" predict embeddings about the same as 7 principal components, corresponding to over 50% of variance in the data. We show that some principal dimensions operate differently for male and female speakers, suggesting there is implicit gender recognition within the embedding systems. However we show that speaker age is not well captured by embeddings, suggesting opportunities exist for improvements in their calculation.