🤖 AI Summary
This study addresses the misalignment between model embedding geometry and human auditory perception caused by relying solely on Equal Error Rate (EER) for voice similarity evaluation. To overcome this limitation, we propose a novel perceptual alignment criterion grounded in embedding geometry. By comparing loss functions such as AAM-Softmax through rank correlation analysis and dimensional bottleneck compression, we reveal that the high-dimensional diffusion induced by classification losses fundamentally contradicts the low-dimensional nature of human perception. Notably, we identify a strong negative correlation of −0.95 between effective dimensionality and perceptual alignment. Applying dimensional bottleneck compression significantly improves the perceptual alignment score from 0.08 to 0.74. These findings provide both a theoretical foundation and a practical framework for voice similarity assessment that is substantially more consistent with human auditory perception.
📝 Abstract
Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). Notably, standard margin-based classification losses (e.g., AAM-Softmax) yield substantially lower perceptual alignment than prototypical metric losses, while EER itself fails to track human judgment, directly challenging the community's implicit assumption. We trace this divergence to embedding geometry, where a model's effective dimensionality ($d_{\mathrm{eff}}$) tracks perceptual alignment with a $-0.95$ rank correlation, revealing that the dimensional spread favored by classification losses fundamentally clashes with the low-dimensional nature of human voice perception. Imposing a dimensionality bottleneck compresses $d_{\mathrm{eff}}$ and raises perceptual alignment ($\rho_{\mathrm{align}}$) from 0.08 to 0.74, establishing a principled geometric criterion for evaluating voice similarity.