🤖 AI Summary
This study addresses the computational challenge of K-means clustering with cardinality constraints on manifolds, where existing methods struggle to balance theoretical guarantees with efficiency. To overcome this, the authors reformulate the constrained problem into an unconstrained Riemannian nonsmooth optimization via a difference-of-convex (DC) penalty, rigorously proving their equivalence and establishing a global error bound. They further propose the RADA-DC algorithm, which synergistically integrates Riemannian geometry, DC programming, and dual regularization techniques for efficient optimization. Theoretically, this work provides an iteration complexity guarantee of O(ε⁻³). Empirically, extensive experiments on large-scale datasets demonstrate that the proposed approach significantly outperforms existing baselines, achieving superior computational efficiency while maintaining high clustering accuracy.
📝 Abstract
K-means is a widely adopted clustering approach in signal processing and machine learning. In this paper, we study K-means clustering through a cardinality-constrained formulation on a compact embedded submanifold. We replace the cardinality constraint with a difference-of-convex (DC) penalty and establish a global error bound to prove that the penalized and constrained formulations share the same global minimizers whenever the penalty parameter exceeds a finite threshold. To solve the resulting nonsmooth Riemannian DC problem, we reformulate it as a minimax problem and propose RADA-DC, a Riemannian alternating descent ascent method combining dual regularization with DC linearization. Under standard assumptions and suitable parameter choices, RADA-DC finds an $ε$-Riemannian critical point within $O(ε^{-3})$ iterations. We conduct experiments on synthetic and real-world datasets to demonstrate that the proposed method outperforms the tested baselines, including K-means++, in solution quality at competitive computational cost when the number of clusters is large.