๐ค AI Summary
This work addresses the concern that using cosine similarity with unnormalized embeddings in matrix factorization models leads to arbitrary results due to gauge freedom, casting doubt on its validity. Through theoretical and geometric analysis, the authors demonstrate that the core issue lies not in cosine similarity itself, but in the mismatch between the training objective and the similarity metric. When embeddings are constrained to the unit hypersphere, gauge freedom is entirely eliminated. Under this constraint, cosine distance becomes strictly equivalent to Euclidean distance: specifically, the cosine distance equals half the squared Euclidean distance, and both induce identical neighbor rankings. This equivalence provides a rigorous theoretical foundation for the principled use of cosine similarity with normalized embeddings.
๐ Abstract
Steck, Ekanadham, and Kallus [arXiv:2403.05440] demonstrate that cosine similarity of learned embeddings from matrix factorization models can be rendered arbitrary by a diagonal ``gauge'' matrix $D$. Their result is correct and important for practitioners who compute cosine similarity on embeddings trained with dot-product objectives. However, we argue that their conclusion, cautioning against cosine similarity in general, conflates the pathology of an incompatible training objective with the geometric validity of cosine distance on the unit sphere. We prove that when embeddings are constrained to the unit sphere $\mathbb{S}^{d-1}$ (either during or after training with an appropriate objective), the $D$-matrix ambiguity vanishes identically, and cosine distance reduces to exactly half the squared Euclidean distance. This monotonic equivalence implies that cosine-based and Euclidean-based neighbor rankings are identical on normalized embeddings. The ``problem'' with cosine similarity is not cosine similarity, it is the failure to normalize.