🤖 AI Summary
Modeling the cosine similarity distribution of sentence embeddings remains challenging due to its bounded support and skewed, multimodal nature.
Method: This work proposes, for the first time, a translated-and-truncated Gamma Mixture Model (GMM) constrained to [−1, 1] to characterize this distribution. Using a fixed corpus, it computes similarity scores between document embeddings and reference query embeddings, then fits the empirical density via expectation-maximization (EM).
Contribution/Results: The model achieves high-fidelity distributional fitting across diverse corpora. Integrating hierarchical topic clustering, we uncover systematic correlations between distributional properties—such as kurtosis and skewness—and semantic topic hierarchies. This provides a novel statistical framework for semantic similarity modeling, enhances interpretability of embedding spaces, and releases an open-source, reusable modeling toolkit to support reproducible research in embedding analysis.
📝 Abstract
We study the cosine similarity of sentence transformer embeddings and observe that they are well modeled by gamma mixtures. From a fixed corpus, we measure similarities between all document embeddings and a reference query embedding. Empirically we find that these distributions are often well captured by a gamma distribution shifted and truncated to [-1,1], and in many cases, by a gamma mixture. We propose a heuristic model in which a hierarchical clustering of topics naturally leads to a gamma-mixture structure in the similarity scores. Finally, we outline an expectation-maximization algorithm for fitting shifted gamma mixtures, which provides a practical tool for modeling similarity distributions.