🤖 AI Summary
To address the challenge of irrelevant features interfering with structural pattern identification in high-dimensional unsupervised clustering, this paper proposes a scalable, stability-driven feature selection framework. Methodologically, it introduces (1) a novel Minkowski-weighted k-means++ initialization strategy that jointly optimizes distance metric learning and cluster center initialization; (2) the FS-MWK++ algorithm and its subsampled variant SFS-MWK++, which enhance scalability via multi-exponent weight aggregation and stochastic subsampling; and (3) a weight-stability-based feature selection paradigm, supported by theoretical convergence guarantees. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach significantly outperforms existing unsupervised feature selection methods, achieving superior trade-offs between clustering accuracy and computational efficiency.
📝 Abstract
Unsupervised feature selection is critical for improving clustering performance in high-dimensional data, where irrelevant features can obscure meaningful structure. In this work, we introduce the Minkowski weighted $k$-means++, a novel initialisation strategy for the Minkowski Weighted $k$-means. Our initialisation selects centroids probabilistically using feature relevance estimates derived from the data itself. Building on this, we propose two new feature selection algorithms, FS-MWK++, which aggregates feature weights across a range of Minkowski exponents to identify stable and informative features, and SFS-MWK++, a scalable variant based on subsampling. We support our approach with a theoretical guarantee under mild assumptions and extensive experiments showing that our methods consistently outperform existing alternatives.