🤖 AI Summary
To address k-means’ limitations—its inability to handle non-convex cluster structures, sensitivity to the pre-specified number of clusters (k), and poor scalability in distributed settings—this paper proposes a lightweight post-clustering optimization framework. Methodologically, it introduces: (1) a geometry-driven, radius-based cluster merging mechanism that hierarchically merges overlapping clusters to recover non-convex shapes and tolerate overestimated (k); and (2) a recursively block-decomposable distributed merging architecture ensuring global consistency while scaling efficiently across large-scale distributed systems. The framework integrates seamlessly into scikit-learn’s k-means implementation without modifying the core algorithm. Evaluated on multiple benchmark datasets, it achieves an average 12.3% improvement in clustering accuracy with less than 5% additional computational overhead, demonstrating high efficiency, robustness, and practical deployability.
📝 Abstract
Traditional k-means clustering underperforms on non-convex shapes and requires the number of clusters k to be specified in advance. We propose a simple geometric enhancement: after standard k-means, each cluster center is assigned a radius (the distance to its farthest assigned point), and clusters whose radii overlap are merged. This post-processing step loosens the requirement for exact k: as long as k is overestimated (but not excessively), the method can often reconstruct non-convex shapes through meaningful merges. We also show that this approach supports recursive partitioning: clustering can be performed independently on tiled regions of the feature space, then globally merged, making the method scalable and suitable for distributed systems. Implemented as a lightweight post-processing step atop scikit-learn's k-means, the algorithm performs well on benchmark datasets, achieving high accuracy with minimal additional computation.