GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance bottlenecks of high GPU memory consumption and frequent data movement in large-scale non-negative low-rank K-means clustering. To overcome these limitations, we propose GEM-KMeans, an algorithm that fuses gradient updates with non-negative projection via a spectrally normalized equivalence formulation, enabling high-bandwidth memory to retain only a single factor matrix. Furthermore, this work introduces a novel IO-aware architecture that unifies gradient computation, non-negativity constraints, and statistics updates as epilogue operations within matrix multiplications, substantially reducing intermediate buffer requirements. Experimental results demonstrate that GEM-KMeans significantly decreases memory overhead while preserving clustering accuracy, exhibiting superior computational efficiency and scalability compared to the conventional Lloyd's algorithm.
📝 Abstract
Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for $K$-means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematically equivalent NLR formulation that fuses the gradient update, nonnegative projection, and sufficient statistics for normalization and iterate movement into a matrix-multiplication epilogue. Instead of retaining three massive factor-sized arrays, our IO-aware GPU implementation materializes only one single factor with small tile-reduction arrays as additional storage in the High Bandwidth Memory (HBM). We derive explicit memory costs and spectrally normalized smoothness bounds for optimizing the clustering objective function. Accurate clustering is demonstrated at massive scales on synthetic and real datasets, where performance gains of GEM-KMeans over existing GPU-accelerated Lloyd's algorithms involve data-dependent runtime tradeoffs.
Problem

Research questions and friction points this paper is trying to address.

Clustering
Memory-efficient
Nonnegative low-rank matrix factorization
GPU optimization
Large-scale data analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Memory-Efficient Clustering
Nonnegative Low-Rank Factorization
GPU Optimization
Matrix-Multiplication Epilogue
Spectral Normalization
🔎 Similar Papers
No similar papers found.