Round-Trip KNN Clustering: multiscale hierarchical cluster detection on directed nearest-neighbour graphs

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses key challenges in multi-scale clustering, including the requirement to predefine cluster numbers, information loss in bidirectional graphs, and erroneous merging caused by sparse bridging. To overcome these limitations, this work proposes RTKNNC, a method that introduces round-trip recursive traversal and weighted voting on directed K-nearest neighbor (KNN) graphs. This approach fully preserves topological relationships to identify hierarchical structures across multiple scales. Furthermore, it integrates an unsupervised Gaussian mixture model to refine subgroups while dynamically detecting cluster persistence and merging behavior. Evaluated independently on benchmark datasets, the proposed algorithm achieves an Adjusted Rand Index exceeding 0.98, outperforming seven mainstream methods. These results demonstrate that RTKNNC enables high-performance, fully unsupervised multi-scale clustering without requiring prior knowledge of cluster counts.
📝 Abstract
We introduce Round-Trip KNN Clustering (RTKNNC), a graph-based method for finding cluster structure at several neighbourhood scales without requiring the number of clusters in advance. Unlike approaches that first make a $k$-nearest-neighbour (KNN) graph undirected, RTKNNC keeps both directions of the neighbour relation: which points a given point selects and which points select it. Incoming selections are treated as weighted votes that help decide which local connections remain visible during a recursive forward-and-reverse traversal. Repeating the procedure for increasing $K$ reveals how groups persist or merge as the neighbourhood scale grows; for the reference inverse-square model before structural refinement, clusters can merge but do not split. Because graph connectivity can occasionally join distinct groups through a sparse bridge or a small region of overlap, we add an optional label-free refinement. It first tests whether an already formed component is better described by two or three Gaussian subpopulations, and accepts a subdivision only when the proposed groups are large enough and consistent with the visible KNN graph. Across eight synthetic datasets and $K=2,\ldots,16$, independent C and Python implementations produced identical partitions in all 120 reference runs. Refinement increased adjusted Rand index from $0.7817$ to $0.9627$ on a variable-density benchmark and from $0.8083$ to $0.9853$ on a sparse-bridge benchmark. Comparisons with seven external clustering methods show competitive performance while preserving a label-free cluster-construction process.
Problem

Research questions and friction points this paper is trying to address.

clustering
directed KNN graph
multiscale hierarchical structure
cluster detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Directed KNN Graph
Multiscale Clustering
Round-Trip Traversal
Label-Free Refinement
Hierarchical Cluster Detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Eraldo Pereira Marinho
Eraldo Pereira Marinho
Professor de Ciência da Computação, Universidade Estadual Paulista UNESP
Análise de aglomeraçõesAprendizado Profundo
Caetano Mazzoni Ranieri
Caetano Mazzoni Ranieri
Sao Paulo State University
Activity recognitionDeep learningInternet of ThingsMachine learningRobotics
F
Fabricio Aparecido Breve
Department of Statistics, Applied Mathematics and Computing, São Paulo State University (UNESP), Institute of Geosciences and Exact Sciences, Av. 24-A, 1515, Rio Claro, São Paulo, Brazil