🤖 AI Summary
This work addresses the lack of theoretical characterization of the relationship between the objective function and feature weights in existing Minkowski weighted k-means algorithms. We provide the first interpretation of the objective function as a power mean aggregation of within-cluster dispersions, revealing how the Minkowski exponent \(p\) governs feature selection behavior. By integrating Minkowski distance, power mean analysis, and optimization theory, we establish an explicit dependency between feature weights and dispersion, proving that the weights follow a power-law distribution. Our analysis offers theoretical guarantees for the suppression of high-dispersion features and derives tight upper and lower bounds on the objective function, structural properties of the weight configuration, and convergence guarantees for the algorithm, thereby unifying the understanding of its underlying mechanics.
📝 Abstract
The Minkowski weighted k-means (mwk-means) algorithm extends classical k-means by incorporating feature weights and a Minkowski distance. Despite its empirical success, its theoretical properties remain insufficiently understood. We show that the mwk-means objective can be expressed as a power-mean aggregation of within-cluster dispersions, with the order determined by the Minkowski exponent p. This formulation reveals how p controls the transition between selective and uniform use of features. Using this representation, we derive bounds for the objective function and characterise the structure of the feature weights, showing that they depend only on relative dispersion and follow a power-law relationship with dispersion ratios. This leads to explicit guarantees on the suppression of high-dispersion features. Finally, we establish convergence of the algorithm and provide a unified theoretical interpretation of its behaviour.