🤖 AI Summary
To address the curse of dimensionality and slow convergence of the Expectation-Maximization (EM) algorithm in Gaussian Mixture Model (GMM)-based clustering of high-dimensional continuous data, this paper proposes a joint embedding-and-clustering optimization framework that enables the first non-sequential, co-optimized integration of Principal Component Analysis (PCA) and Classification EM (CEM). Unlike conventional two-stage pipelines, our method simultaneously optimizes both the low-dimensional embedding space and GMM parameters, eliminating error propagation. Theoretically, it unifies PCA, K-means, CEM, and spectral clustering under a single coherent formulation. Empirical evaluation demonstrates that the proposed approach accelerates convergence by 2–5× over standard EM, improves clustering accuracy, and yields more interpretable embeddings. It consistently outperforms state-of-the-art baselines on multiple high-dimensional benchmark datasets.
📝 Abstract
The mixture model is undoubtedly one of the greatest contributions to clustering. For continuous data, Gaussian models are often used and the Expectation-Maximization (EM) algorithm is particularly suitable for estimating parameters from which clustering is inferred. If these models are particularly popular in various domains including image clustering, they however suffer from the dimensionality and also from the slowness of convergence of the EM algorithm. However, the Classification EM (CEM) algorithm, a classifying version, offers a fast convergence solution while dimensionality reduction still remains a challenge. Thus we propose in this paper an algorithm combining simultaneously and non-sequentially the two tasks --Data embedding and Clustering-- relying on Principal Component Analysis (PCA) and CEM. We demonstrate the interest of such approach in terms of clustering and data embedding. We also establish different connections with other clustering approaches.