🤖 AI Summary
This study addresses key challenges in multi-scale clustering, including the requirement to predefine cluster numbers, information loss in bidirectional graphs, and erroneous merging caused by sparse bridging. To overcome these limitations, this work proposes RTKNNC, a method that introduces round-trip recursive traversal and weighted voting on directed K-nearest neighbor (KNN) graphs. This approach fully preserves topological relationships to identify hierarchical structures across multiple scales. Furthermore, it integrates an unsupervised Gaussian mixture model to refine subgroups while dynamically detecting cluster persistence and merging behavior. Evaluated independently on benchmark datasets, the proposed algorithm achieves an Adjusted Rand Index exceeding 0.98, outperforming seven mainstream methods. These results demonstrate that RTKNNC enables high-performance, fully unsupervised multi-scale clustering without requiring prior knowledge of cluster counts.
📝 Abstract
We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help decide which local connections remain visible during a recursive forward-and-reverse traversal. Repeating the procedure for increasing $K$ reveals how groups persist or merge as the neighbourhood scale grows; for the reference inverse-square model before structural refinement, clusters can merge but do not split. Because graph connectivity can occasionally join distinct groups through a sparse bridge or a small region of overlap, we add an optional label-free refinement. It first tests whether an already formed component is better described by two or three Gaussian subpopulations, and accepts a subdivision only when the proposed groups are large enough and consistent with the visible KNN graph. Across eight synthetic datasets and $K=2,\ldots,16$, independent C and Python implementations produced identical partitions in all 120 reference runs. Refinement increased adjusted Rand index from $0.7817$ to $0.9627$ on a variable-density benchmark and from $0.8083$ to $0.9853$ on a sparse-bridge benchmark. Comparisons with seven external clustering methods show competitive performance while preserving a label-free cluster-construction process.