Score
Design and implement post-processing procedures that refine model outputs by retrieving k nearest neighbors in a chosen feature or input space and aggregating their labels or predictions to adjust final labels, confidence scores, or structured outputs; this includes choosing similarity metrics and k, implementing neighbor-aggregation rules (e.g., majority vote, distance-weighted averaging, label smoothing), enforcing constraints such as hierarchical consistency, and measuring effects on accuracy, calibration, and consistency.
To address k-means’ limitations—its inability to handle non-convex cluster structures, sensitivity to the pre-specified number of clusters (k), and poor scalability in distributed settings—this paper proposes a lightweight post-clustering optimization framework. Methodologically, it introduces: (1) a geometry-driven, radius-based cluster merging mechanism that hierarchically merges overlapping clusters to recover non-convex shapes and tolerate overestimated (k); and (2) a recursively block-decomposable distributed merging architecture ensuring global consistency while scaling efficiently across large-scale distributed systems. The framework integrates seamlessly into scikit-learn’s k-means implementation without modifying the core algorithm. Evaluated on multiple benchmark datasets, it achieves an average 12.3% improvement in clustering accuracy with less than 5% additional computational overhead, demonstrating high efficiency, robustness, and practical deployability.
To address the high computational cost, low accuracy, and poor generalizability of distance computation in high-dimensional approximate k-nearest neighbor (AKNN) search, this paper proposes a data-distribution-aware orthogonal projection distance estimation method with a decoupled, data-driven correction scheme. It is the first to incorporate explicit data distribution modeling into orthogonal projection-based distance estimation and fully decouples distance approximation from correction—thereby jointly optimizing efficiency, accuracy, and generality. The approach integrates orthogonal projection for dimensionality reduction, a lightweight data-driven correction model, high-dimensional index optimization, and accelerated distance computation mechanisms. Extensive experiments on multiple real-world datasets demonstrate that our method achieves 1.6–2.1× higher retrieval speed than ADSampling, while significantly improving recall and distance estimation accuracy.
Despite widespread empirical success, the theoretical foundations and practical deployment guidelines for k-nearest neighbors (k-NN) in predictive tasks remain inadequately understood. Method: We establish the first non-asymptotic error bound framework tailored for real-world deployment, replacing conventional smoothness or margin assumptions with verifiable cluster structure as the key success criterion. We integrate approximate nearest neighbor techniques—including LSH and graph-based indexing—and unify k-NN theory with emerging paradigms such as random forests, graphon modeling, and crowdsourcing. We further introduce a novel distance-learning perspective, characterizing how ensemble methods implicitly learn neighborhood structure. Contribution/Results: Evaluated on time-series forecasting, recommender systems, and medical image segmentation, our framework demonstrates that high accuracy is guaranteed solely under cluster-structured data—enhancing both theoretical interpretability and engineering tractability. It provides actionable, error-tolerance-driven guidance for data volume and hyperparameter selection.
This study addresses the lack of theoretical guarantees for $k$-nearest neighbors (kNN) regression when applied to survey data arising from complex sampling designs, which violate the standard i.i.d. assumption. The paper establishes the first consistency framework for kNN regression under such settings by integrating probability sampling theory, nonparametric regression, and asymptotic analysis. It rigorously derives a lower bound on the convergence rate and demonstrates that kNN regression remains consistent even when accounting for sampling weights and dependence structures inherent in complex surveys. However, the convergence rate is shown to be adversely affected by the curse of dimensionality. These theoretical findings extend the classical kNN consistency results—previously limited to i.i.d. data—to realistic survey contexts. The validity of the theoretical conclusions is corroborated through both simulation studies and analyses of real-world survey data.
Gao et al. (JASA 2022) proposed a post-clustering inference framework limited to i.i.d. Gaussian data, failing to accommodate arbitrary dependency structures among observations and features. Method: We generalize their framework to arbitrary dependence settings by developing a unified, dependence-aware post-clustering inference methodology compatible with hierarchical agglomerative clustering (single/complete/average linkage) and k-means. We derive theoretical conditions for well-defined p-values that ensure selective Type I error control and enable consistent covariance matrix estimation. Contribution/Results: Integrating selective inference, high-dimensional statistics, and covariance structure modeling, we design a robust testing pipeline. Experiments on synthetic and real-world protein structural data demonstrate substantial improvements in statistical reliability and practical utility for testing mean differences between clusters under dependence.
This study addresses the common trade-off in post-processing ensemble forecasts, where neural network–based calibration often sacrifices predictive sharpness—particularly for short lead times—to improve reliability. To jointly optimize both calibration and sharpness, the authors propose a novel approach that introduces, for the first time, a sharpness penalty term directly into the continuous ranked probability score (CRPS) loss function. Assuming a Gaussian predictive distribution, the method is evaluated on ECMWF 2-meter temperature ensemble forecasts and achieves a reduction of 8.2%–12.5% in the width of central prediction intervals while maintaining CRPS and mean RMSE at levels comparable to baseline methods. This demonstrates a significant enhancement in forecast sharpness without compromising overall forecasting skill.
This work investigates the interplay between interpolation and aggregation in regression tasks and its implications for learnability. By introducing the γ-graph dimension, the study characterizes the learnability boundary for a broad class of natural aggregation procedures and proposes a minimalist aggregation method that takes the median of three interpolating hypotheses. Theoretical analysis demonstrates that this median aggregation achieves optimal sample complexity among all finite interpolating aggregations and strictly outperforms standard interpolating learning. Moreover, the work reveals that certain hypothesis classes are learnable only via infinite or non-interpolating aggregations, thereby establishing fundamental limitations and optimality conditions for finite interpolating aggregation schemes.
This work addresses the problem of aggregating multiple calibrated Bayesian expert forecasts to construct a new predictor that remains calibrated and is Blackwell-dominated by the target expert, rather than merely minimizing a specific loss function. Under the setting where only the experts’ prior distributions are observed—without access to the true state—the authors formally define the aggregation objective as simultaneously achieving calibration and Blackwell refinability. By modeling calibrated experts through reduced-form information structures, they characterize the set of feasible predictions using the row space of a linear system intersected with a non-negative cone, and analyze it via Blackwell dominance theory. Their main contributions include efficient solvability of both randomized aggregation problems, while showing that determining the existence of a deterministic aggregator is NP-hard and admits no multiplicative PTAS unless P = NP, thereby revealing a fundamental computational distinction between randomized and deterministic aggregation.
This work proposes a general, model-agnostic post-processing framework that systematically integrates ensemble learning into fairness optimization. By aggregating predictions from multiple base models without requiring access to their internal architectures, the approach uniformly supports diverse predictive tasks—including classification, regression, and survival analysis—and accommodates a wide range of fairness definitions. The method addresses the inherent trade-off between predictive performance and fairness in machine learning models. Experimental results demonstrate that it significantly enhances fairness while maintaining or only marginally compromising prediction accuracy, thereby validating its effectiveness and generalizability across multiple scenarios.