Shapley-Inspired Feature Weighting in $k$-means with No Additional Hyperparameters

📅 2025-08-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In high-dimensional or noisy data, conventional clustering methods assume uniform feature contribution, often degrading performance; existing feature-weighting approaches further rely on additional hyperparameters. This paper proposes SHARK—a hyperparameter-free, Shapley-value-inspired feature-weighted k-means algorithm. Methodologically, SHARK is the first to axiomatically integrate Shapley values into unsupervised clustering by decomposing the k-means objective function to quantify per-feature importance. It introduces a polynomial-time approximation algorithm to circumvent exponential computational complexity and designs an iterative reweighting scheme for adaptive feature selection. Evaluated on synthetic and real-world benchmarks, SHARK consistently achieves superior clustering accuracy and robustness—particularly under noise—outperforming state-of-the-art baseline methods.

Technology Category

Machine Learning: ClusteringData Mining & Knowledge Management: Anomaly/Outlier DetectionReasoning under Uncertainty: Stochastic Optimization

Application Category

Web Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
Clustering algorithms often assume all features contribute equally to the data structure, an assumption that usually fails in high-dimensional or noisy settings. Feature weighting methods can address this, but most require additional parameter tuning. We propose SHARK (Shapley Reweighted $k$-means), a feature-weighted clustering algorithm motivated by the use of Shapley values from cooperative game theory to quantify feature relevance, which requires no additional parameters beyond those in $k$-means. We prove that the $k$-means objective can be decomposed into a sum of per-feature Shapley values, providing an axiomatic foundation for unsupervised feature relevance and reducing Shapley computation from exponential to polynomial time. SHARK iteratively re-weights features by the inverse of their Shapley contribution, emphasising informative dimensions and down-weighting irrelevant ones. Experiments on synthetic and real-world data sets show that SHARK consistently matches or outperforms existing methods, achieving superior robustness and accuracy, particularly in scenarios where noise may be present. Software: https://github.com/rickfawley/shark.
Problem

Research questions and friction points this paper is trying to address.

Overcoming equal feature weight assumption in clustering
Avoiding extra hyperparameters in feature weighting
Enhancing robustness and accuracy in noisy data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Shapley values quantify feature relevance
No additional hyperparameters beyond k-means
Polynomial time Shapley computation
🔎 Similar Papers