🤖 AI Summary
This study addresses the challenge of accurately estimating the number of clusters in datasets with unknown clustering structure by proposing a nonparametric method based on pairwise distances. The approach generates p-values through dependence-adjusted multiple hypothesis testing that adapts to sample size and employs a stepwise selection strategy to infer the optimal number of clusters, without requiring a pre-specified cluster count. It is applicable to data of arbitrary dimensionality and compatible with existing clustering algorithms. Experimental results demonstrate that the method substantially reduces computational overhead while achieving superior accuracy and robustness in cluster number estimation compared to current state-of-the-art metrics, offering strong theoretical grounding and practical utility.
📝 Abstract
This paper proposes a novel, nonparametric, interpoint distance-based measure to investigate whether there exist any groups in a set of given data, and if so then, how many groups are prevailing in total. It is a cluster accuracy index useful for arbitrary-dimensional data set, in association with any clustering algorithm having the number of groups specified as a priori. We perform univariate, nonparametric, multiple statistical tests of hypotheses, where as many dependent tests as the sample size are carried out using the interpoint distances. They possess $p$-values to be combined to reach a decision, which is taken in a step-wise process for a possible number of clusters. It reduces the unnecessary computations compared with the other accuracy measures from the literature. Data study establishes the proposed index's efficiency and superiority.