Improving internal cluster quality evaluation in noisy Gaussian mixtures

📅 2025-03-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Internal clustering validation metrics—such as silhouette coefficient, Calinski-Harabasz index, and Davies-Bouldin index—are prone to distortion in high-dimensional noisy data due to interference from irrelevant or redundant features. To address this, we propose Feature Importance Rescaling (FIR), the first method to integrate a dynamic rescaling mechanism—based on feature dispersion—into internal validation. FIR adaptively attenuates the influence of noisy or redundant features on compactness and separation measurements, without requiring ground-truth labels and while remaining compatible with mainstream indices. Evaluated on multiple noisy Gaussian mixture datasets, FIR significantly enhances robustness: average Spearman correlation with true cluster structures improves by 37%, estimation variance decreases by 52%, and discriminative capability remains stable even under high cluster overlap. FIR thus improves both the accuracy and interpretability of internal validation in high-dimensional settings.

Technology Category

Machine Learning: Dimensionality Reduction/Feature SelectionData Mining & Knowledge Management: Anomaly/Outlier DetectionSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating success
📝 Abstract
Clustering is a fundamental technique in machine learning and data analysis, widely used across various domains. Internal clustering validation measures, such as the Average Silhouette Width, Calinski-Harabasz, and Davies-Bouldin indices, play a crucial role in assessing clustering quality when external ground truth labels are unavailable. However, these measures can be affected by feature relevance, potentially leading to unreliable evaluations in high-dimensional or noisy data sets. In this paper, we introduce a Feature Importance Rescaling (FIR) method designed to enhance internal clustering validation by adjusting feature contributions based on their dispersion. Our method systematically attenuates noise features making clustering compactness and separation clearer, and by consequence aligning internal validation measures more closely with the ground truth. Through extensive experiments on synthetic data sets under different configurations, we demonstrate that FIR consistently improves the correlation between internal validation indices and the ground truth, particularly in settings with noisy or irrelevant features. The results show that FIR increases the robustness of clustering evaluation, reduces variability in performance across different data sets, and remains effective even when clusters exhibit significant overlap. These findings highlight the potential of FIR as a valuable enhancement for internal clustering validation, making it a practical tool for unsupervised learning tasks where labelled data is not available.
Problem

Research questions and friction points this paper is trying to address.

Enhances internal clustering validation in noisy data
Adjusts feature contributions to improve clustering quality
Increases robustness and reduces variability in clustering evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feature Importance Rescaling (FIR) enhances clustering validation
FIR adjusts feature contributions based on dispersion
FIR reduces noise impact, improving clustering evaluation robustness
🔎 Similar Papers
No similar papers found.