Evaluation of the number of clusters in a data set using $p$-values from Multiple Tests of Hypotheses

📅 2026-05-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of accurately estimating the number of clusters in datasets with unknown clustering structure by proposing a nonparametric method based on pairwise distances. The approach generates p-values through dependence-adjusted multiple hypothesis testing that adapts to sample size and employs a stepwise selection strategy to infer the optimal number of clusters, without requiring a pre-specified cluster count. It is applicable to data of arbitrary dimensionality and compatible with existing clustering algorithms. Experimental results demonstrate that the method substantially reduces computational overhead while achieving superior accuracy and robustness in cluster number estimation compared to current state-of-the-art metrics, offering strong theoretical grounding and practical utility.
📝 Abstract
This paper proposes a novel, nonparametric, interpoint distance-based measure to investigate whether there exist any groups in a set of given data, and if so then, how many groups are prevailing in total. It is a cluster accuracy index useful for arbitrary-dimensional data set, in association with any clustering algorithm having the number of groups specified as a priori. We perform univariate, nonparametric, multiple statistical tests of hypotheses, where as many dependent tests as the sample size are carried out using the interpoint distances. They possess $p$-values to be combined to reach a decision, which is taken in a step-wise process for a possible number of clusters. It reduces the unnecessary computations compared with the other accuracy measures from the literature. Data study establishes the proposed index's efficiency and superiority.
Problem

Research questions and friction points this paper is trying to address.

cluster number estimation
interpoint distance
nonparametric test
multiple hypothesis testing
p-value combination
Innovation

Methods, ideas, or system contributions that make the work stand out.

nonparametric
interpoint distance
multiple hypothesis testing
cluster number estimation
p-value combination
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Soumita Modak
Faculty, Department of Statistics, University of Calcutta, Basanti Devi College, 147B, Rash Behari Ave, Kolkata-700029, India