Scalable unsupervised feature selection via weight stability

📅 2025-06-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of irrelevant features interfering with structural pattern identification in high-dimensional unsupervised clustering, this paper proposes a scalable, stability-driven feature selection framework. Methodologically, it introduces (1) a novel Minkowski-weighted k-means++ initialization strategy that jointly optimizes distance metric learning and cluster center initialization; (2) the FS-MWK++ algorithm and its subsampled variant SFS-MWK++, which enhance scalability via multi-exponent weight aggregation and stochastic subsampling; and (3) a weight-stability-based feature selection paradigm, supported by theoretical convergence guarantees. Extensive experiments on multiple benchmark datasets demonstrate that the proposed approach significantly outperforms existing unsupervised feature selection methods, achieving superior trade-offs between clustering accuracy and computational efficiency.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionData Mining & Knowledge Management: Scalability, Parallel & Distributed SystemsReasoning under Uncertainty: Stochastic Optimization

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search engines
📝 Abstract
Unsupervised feature selection is critical for improving clustering performance in high-dimensional data, where irrelevant features can obscure meaningful structure. In this work, we introduce the Minkowski weighted $k$-means++, a novel initialisation strategy for the Minkowski Weighted $k$-means. Our initialisation selects centroids probabilistically using feature relevance estimates derived from the data itself. Building on this, we propose two new feature selection algorithms, FS-MWK++, which aggregates feature weights across a range of Minkowski exponents to identify stable and informative features, and SFS-MWK++, a scalable variant based on subsampling. We support our approach with a theoretical guarantee under mild assumptions and extensive experiments showing that our methods consistently outperform existing alternatives.
Problem

Research questions and friction points this paper is trying to address.

Improve clustering in high-dimensional data by selecting relevant features
Develop scalable unsupervised feature selection algorithms
Enhance feature weight stability across Minkowski exponents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Minkowski weighted k-means++ initialization strategy
FS-MWK++ aggregates feature weights for stability
SFS-MWK++ enables scalable subsampling variant
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xudong Zhang
School of Computer Science and Electronic Engineering, University of Essex, Wivenhoe, UK.
Renato Cordeiro de Amorim
Renato Cordeiro de Amorim
Senior Lecturer in Computer Science and AI at the University of Essex
ClusteringFeature SelectionUnsupervised LearningCombinatorics