Score
Designs and implements classifiers that assign labels by finding nearest neighbors in a space whose elements are rankings (ranked lists or permutations), using rank-correlation measures as the distance metric. Builds algorithms to compute ranking correlations, identify nearest neighbors among rankings, assign class labels from closest rankings, and evaluate how well ranking distances separate classes.
This work addresses the poor interpretability and rigid decision boundaries of nearest-neighbor (NN) classification. We propose the first fully category-theoretic reconstruction framework for NN methods. Methodologically, we model data and label spaces as Lawvere metric spaces, extend the original dataset via profunctors, and *purely deductively derive* both 1-NN and k-NN classifiers within this categorical structure. This derivation naturally yields weighted Voronoi tessellations and introduces a label metric to define soft, geometry-aware decision boundaries—enabling continuous, label-structure-dependent predictions. Key contributions include: (i) establishing the first category-theory-driven, interpretable theoretical foundation for NN classification; (ii) unifying 1-NN and k-NN under a single formal categorical framework; and (iii) enabling flexible, label-metric-guided generalization, thereby substantially enhancing model transparency and adaptability. The framework bridges abstract category theory with practical machine learning, offering principled insights into NN behavior beyond heuristic justification.
To address the limited generalization capability of k-nearest neighbors (k-NN) ensemble methods, this paper proposes an adaptive k-NN classifier based on discriminative subspace projection, embedded within the Bootstrap aggregating (bagging) framework. For each base classifier, trained on a bootstrap sample, we jointly learn an optimal discriminative projection direction and adaptively select the neighborhood size k—thereby simultaneously enhancing discriminability and ensemble diversity. Extensive experiments across multiple benchmark datasets demonstrate that the proposed ensemble significantly outperforms random forests and state-of-the-art k-NN ensembles. An open-source R package ensures reproducibility and practical applicability. The key contributions are: (1) the first k-NN ensemble framework integrating discriminative subspace learning with adaptive k-selection; and (2) a systematic improvement in the generalization performance of k-NN within bagging, achieved by explicitly increasing both the effectiveness and dissimilarity of base classifiers.
Traditional clustering methods emphasize stability of cluster centers, neglecting the practical need for stability of point-level labels—the named identifiers indicating each sample’s cluster assignment. Method: This paper formally defines “label consistency” as the pointwise label distance between successive clustering solutions, departing from the conventional center-stability paradigm. For the $k$-center and $k$-median problems, we propose the first theoretically grounded consistency-aware approximation algorithm. Leveraging combinatorial optimization and metric-space analysis, we design a dynamic label-distance modeling framework coupled with constrained optimization to jointly optimize clustering quality and label stability. Contribution/Results: Our algorithm achieves an $O(1)$-approximation ratio for both objectives and provides a tight upper bound on the label change rate. It establishes an optimal trade-off between clustering accuracy and label consistency, offering a novel, interpretable, and deployable paradigm for dynamic clustering.
Despite widespread empirical success, the theoretical foundations and practical deployment guidelines for k-nearest neighbors (k-NN) in predictive tasks remain inadequately understood. Method: We establish the first non-asymptotic error bound framework tailored for real-world deployment, replacing conventional smoothness or margin assumptions with verifiable cluster structure as the key success criterion. We integrate approximate nearest neighbor techniques—including LSH and graph-based indexing—and unify k-NN theory with emerging paradigms such as random forests, graphon modeling, and crowdsourcing. We further introduce a novel distance-learning perspective, characterizing how ensemble methods implicitly learn neighborhood structure. Contribution/Results: Evaluated on time-series forecasting, recommender systems, and medical image segmentation, our framework demonstrates that high accuracy is guaranteed solely under cluster-structured data—enhancing both theoretical interpretability and engineering tractability. It provides actionable, error-tolerance-driven guidance for data volume and hyperparameter selection.
To address the high computational cost of k-nearest neighbors (kNN) in multi-label classification—stemming from large-scale training sets—this paper pioneers the extension of prototype learning to the multi-label setting. We propose a label-aware prototype generation method that jointly optimizes label structure consistency and instance similarity. Our approach integrates a multi-label distance metric, greedy initialization, iterative optimization guided by label coverage, and an adaptive kNN reweighting mechanism. Experiments across multiple benchmark datasets demonstrate that our method compresses the training set by over 80%, while maintaining or improving macro-F1 score and classification accuracy. Crucially, it significantly reduces inference cost without sacrificing performance. The core contribution is the first interpretable and efficient prototype learning framework specifically designed for multi-label classification, bridging scalability and fidelity in label-space modeling.
The original construction of metric skip lists exhibits inherent sequentiality, which hinders parallelization and limits their efficiency in large-scale nearest neighbor search. This work proposes the first work-efficient, polylogarithmic-span parallel construction algorithm for metric skip lists. Relying only on a constant expansion rate—without requiring a bounded aspect ratio—and leveraging a divide-and-conquer strategy combined with randomized analysis, the algorithm achieves an expected $O(n \log n)$ total work and polylogarithmic depth with high probability. The method supports nearest neighbor search and several downstream applications, including bichromatic closest pair, density-based clustering, and k-nearest neighbor graph construction, offering the first solution that simultaneously guarantees both work efficiency and low parallel depth for these tasks.
This study addresses the meta-classification problem of one-class classification (OCC) models, aiming to identify their corresponding training datasets, algorithms, or hyperparameters. To this end, the work proposes a unified representation of OCC models as normality rankings and reframes the meta-classification task as a ranking-based dataset classification problem. The authors develop a cohesive framework that integrates nearest-neighbor classifiers with rank correlation measures to jointly classify models, data, and rankings. Experimental results demonstrate that the proposed approach achieves high accuracy in dataset label classification and effectively distinguishes between OCC models trained with different algorithms within the same category. To the best of the authors’ knowledge, this is the first method to enable systematic identification of meta-information associated with OCC models.
This work addresses the design of nearest neighbor search data structures tailored to a given query distribution. It proposes the first algorithm capable of efficiently learning an approximately optimal balanced halfspace partitioning tree under Gaussian-like distributional assumptions. By formulating tree construction as a balanced halfspace cut problem and integrating polynomial threshold functions with a distribution-aware learning strategy, the method circumvents the NP-hard regularized optimization typically involved. Under the assumption that a perfect partitioning tree exists, the approach achieves query time better than $O(nd)$ while ensuring provably bounded cutting error in the learned tree, thereby significantly enhancing both the efficiency and theoretical guarantees of data-driven nearest neighbor search.
This work addresses the longstanding challenge of constructing lower bounds for sign rank by establishing a novel connection between the ℤ₂-index and list replicability. We prove, for the first time, that the ℤ₂-index is linearly upper-bounded by the list replicability number, thereby achieving a strong separation between sign rank and the ℤ₂-index. This insight yields new combinatorial upper bounds and a composition lemma for list replicability. By integrating tools from combinatorics, topological methods, and representation complexity, we resolve an open problem posed by Frick et al., uncovering intrinsic relationships between list replicability and combinatorial parameters such as height and minimum star number. Furthermore, we provide an upper bound on the list replicability of product concept classes.
This study addresses the challenge of effectively classifying graphs with heterogeneous structures and sizes without relying on node embeddings or graph alignment. To this end, the authors extend the k-nearest neighbors (kNN) classifier to metric spaces induced by the Gromov–Wasserstein (GW) and fused Gromov–Wasserstein (fGW) distances, treating graphs as finite metric measure spaces equipped with node features. The work establishes, for the first time, the universal consistency of both GW-kNN and fGW-kNN classifiers over the space of graphs and their attribute-augmented extensions, thereby providing a rigorous theoretical foundation for nonparametric classification based on GW-type distances. Empirical evaluations demonstrate that the proposed approach achieves strong performance and excellent generalization across multiple graph benchmark datasets.