Revisiting Label-Free Speaker Embedding Enhancement with vMF Profile Likelihood

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that overly structured modeling in unsupervised speaker embedding enhancement leads to complex and unstable optimization. To mitigate this, we propose a concise closed-form solution based on the von Mises-Fisher (vMF) distribution. By modeling clean targets via vMF likelihood and marginalizing out the concentration parameter, an adaptively weighted optimization objective is constructed, effectively circumventing the redundancy inherent in highly structured formulations. This approach significantly simplifies the enhancement pipeline while improving robustness. Evaluated across multiple datasets including VoxCeleb1, the proposed method not only preserves baseline performance but also achieves substantial gains under mismatched conditions, consistently outperforming diffusion-based baselines overall.
📝 Abstract
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
Problem

Research questions and friction points this paper is trying to address.

speaker embedding enhancement
label-free
speaker verification
acoustic mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speaker Embedding Enhancement
Label-Free Learning
von Mises-Fisher Distribution
Profile Likelihood
Speaker Verification
Seunghwan Kim
Seunghwan Kim
Seoul National University
J
Jinyong Kim
Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea
S
Sooyoung Yang
Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea
Y
Youngjin Ko
Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea
M
Myungjoo Kang
1 Interdisciplinary Program in Artificial Intelligence, Seoul National University, South Korea; 2 Department of Mathematical Sciences, Seoul National University, South Korea; 3 Research Institute of Mathematics, South Korea