๐ค AI Summary
Existing boundary point detection methods are sensitive to density heterogeneity and struggle to identify boundary points in concave structures and high-dimensional manifolds, thereby limiting downstream clustering and classification performance. To address this, we propose LoDD (Local Directional Dispersion), a density-agnostic boundary criterion. LoDD innovatively quantifies local directional dispersion by analyzing eigenvalues of the KNN neighborhoodโs covariance matrix to characterize directional uniformity. It further introduces a gridded distribution assumption to enable adaptive estimation of critical parameters. The method integrates density-agnostic KNN search, directional feature modeling, and structure-driven parameter optimization. Extensive experiments on synthetic and real-world datasets, deep learning training set partitioning, and point cloud void detection demonstrate that LoDD significantly outperforms state-of-the-art approaches. Code and benchmark datasets are publicly available.
๐ Abstract
Boundary point detection aims to outline the external contour structure of clusters and enhance the inter-cluster discrimination, thus bolstering the performance of the downstream classification and clustering tasks. However, existing boundary point detectors are sensitive to density heterogeneity or cannot identify boundary points in concave structures and high-dimensional manifolds. In this work, we propose a robust and efficient boundary point detection method based on Local Direction Dispersion (LoDD). The core of boundary point detection lies in measuring the difference between boundary points and internal points. It is a common observation that an internal point is surrounded by its neighbors in all directions, while the neighbors of a boundary point tend to be distributed only in a certain directional range. By considering this observation, we adopt density-independent K-Nearest Neighbors (KNN) method to determine neighboring points and design a centrality metric LoDD using the eigenvalues of the covariance matrix to depict the distribution uniformity of KNN. We also develop a grid-structure assumption of data distribution to determine the parameters adaptively. The effectiveness of LoDD is demonstrated on synthetic datasets, real-world benchmarks, and application of training set split for deep learning model and hole detection on point cloud data. The datasets and toolkit are available at: https://github.com/ZPGuiGroupWhu/lodd.