🤖 AI Summary
This study addresses the limitations of supervised learning in speaker recognition, which relies heavily on labeled data and exhibits poor generalization—especially under unknown conditions. The work presents a systematic investigation of self-supervised learning (SSL) for this task, offering the first comprehensive evaluation within a unified experimental framework of instance-invariance-based methods such as SimCLR, MoCo, and DINO. The analysis covers key architectural components, hyperparameter sensitivity, and in-domain/out-of-domain generalization capabilities. Results show that DINO achieves the best downstream performance by effectively modeling intra-speaker variability, albeit with high sensitivity to hyperparameters. In contrast, SimCLR and MoCo demonstrate greater robustness, excelling at capturing inter-speaker variability while avoiding representation collapse. This work elucidates the mechanistic differences among SSL approaches in modeling speaker variability and establishes a systematic benchmark and practical guidance for unsupervised speaker recognition.