A Graph-based Approach to Estimating the Number of Clusters

📅 2024-02-23
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenging problem of automatically estimating the number of clusters $k$ in high-dimensional data. We propose a nonparametric method based on similarity graphs: it constructs a robust graph structure, extracts intrinsic similarity patterns via spectral analysis of the graph Laplacian, and introduces an analytically tractable graph statistic—marking the first approach to directly leverage graph-structural features for $k$ selection. We establish theoretical consistency of the estimator under high-dimensional asymptotics. Unlike existing methods, our framework imposes no assumptions on underlying data distributions or specific clustering algorithms, and is both dimensionality-agnostic and computationally efficient. Extensive experiments on synthetic benchmarks, medical imaging, and RNA-seq datasets demonstrate substantial improvements in estimation accuracy under high-dimensional settings, along with enhanced robustness and generalizability.

Technology Category

Machine Learning: ClusteringData Mining & Knowledge Management: Graph Mining, Social Network Analysis & CommunityReasoning under Uncertainty: Graphical Models

Application Category

Graph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsWeb Mining and Content Analysis: Robustness and generalizability of Web mining methodsSemantics and Knowledge: Scalable techniques for the creation, curation, publication, maintenance, and consumption of large, Web-based, structured, reusable, knowledge graphs and ontologies
📝 Abstract
We consider the problem of estimating the number of clusters (k) in a dataset. We propose a non-parametric approach to the problem that utilizes similarity graphs to construct a robust statistic that effectively captures similarity information among observations. This graph-based statistic is applicable to datasets of any dimension, is computationally efficient to obtain, and can be paired with any kind of clustering technique. Asymptotic theory is developed to establish the selection consistency of the proposed approach. Simulation studies demonstrate that the graph-based statistic outperforms existing methods for estimating k, especially in the high-dimensional setting. We illustrate its utility on an imaging dataset and an RNA-seq dataset.
Problem

Research questions and friction points this paper is trying to address.

Estimating cluster count in high-dimensional datasets
Developing graph-based non-parametric clustering method
Improving accuracy in high-dimensional cluster estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-based non-parametric cluster estimation
Robust similarity graph statistic
Dimension-independent computational efficiency
🔎 Similar Papers
No similar papers found.
Iowa State University
Y
Yichuan Bai
Department of Statistics, Iowa State University
L
Lynna Chu
Department of Statistics, Iowa State University