Gamma Mixture Modeling for Cosine Similarity in Small Language Models

📅 2025-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Modeling the cosine similarity distribution of sentence embeddings remains challenging due to its bounded support and skewed, multimodal nature. Method: This work proposes, for the first time, a translated-and-truncated Gamma Mixture Model (GMM) constrained to [−1, 1] to characterize this distribution. Using a fixed corpus, it computes similarity scores between document embeddings and reference query embeddings, then fits the empirical density via expectation-maximization (EM). Contribution/Results: The model achieves high-fidelity distributional fitting across diverse corpora. Integrating hierarchical topic clustering, we uncover systematic correlations between distributional properties—such as kurtosis and skewness—and semantic topic hierarchies. This provides a novel statistical framework for semantic similarity modeling, enhances interpretability of embedding spaces, and releases an open-source, reusable modeling toolkit to support reproducible research in embedding analysis.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Large Multimodal Models (LMMs)Reasoning under Uncertainty: Relational Probabilistic Models

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systems
📝 Abstract
We study the cosine similarity of sentence transformer embeddings and observe that they are well modeled by gamma mixtures. From a fixed corpus, we measure similarities between all document embeddings and a reference query embedding. Empirically we find that these distributions are often well captured by a gamma distribution shifted and truncated to [-1,1], and in many cases, by a gamma mixture. We propose a heuristic model in which a hierarchical clustering of topics naturally leads to a gamma-mixture structure in the similarity scores. Finally, we outline an expectation-maximization algorithm for fitting shifted gamma mixtures, which provides a practical tool for modeling similarity distributions.
Problem

Research questions and friction points this paper is trying to address.

Modeling cosine similarity distributions using gamma mixtures
Analyzing sentence transformer embeddings in small language models
Developing expectation-maximization algorithm for shifted gamma mixtures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gamma mixture modeling for cosine similarity
Heuristic model with hierarchical topic clustering
Expectation-maximization algorithm for fitting distributions
🔎 Similar Papers
No similar papers found.
K
Kevin Player
Software Engineering Institute, Carnegie Mellon University