knn post-processing

Design and implement post-processing procedures that refine model outputs by retrieving k nearest neighbors in a chosen feature or input space and aggregating their labels or predictions to adjust final labels, confidence scores, or structured outputs; this includes choosing similarity metrics and k, implementing neighbor-aggregation rules (e.g., majority vote, distance-weighted averaging, label smoothing), enforcing constraints such as hierarchical consistency, and measuring effects on accuracy, calibration, and consistency.

knnpost-processing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address k-means’ limitations—its inability to handle non-convex cluster structures, sensitivity to the pre-specified number of clusters (k), and poor scalability in distributed settings—this paper proposes a lightweight post-clustering optimization framework. Methodologically, it introduces: (1) a geometry-driven, radius-based cluster merging mechanism that hierarchically merges overlapping clusters to recover non-convex shapes and tolerate overestimated (k); and (2) a recursively block-decomposable distributed merging architecture ensuring global consistency while scaling efficiently across large-scale distributed systems. The framework integrates seamlessly into scikit-learn’s k-means implementation without modifying the core algorithm. Evaluated on multiple benchmark datasets, it achieves an average 12.3% improvement in clustering accuracy with less than 5% additional computational overhead, demonstrating high efficiency, robustness, and practical deployability.

Enables scalable distributed clustering through recursive partitioningEnhances k-means for non-convex shapes via radius-guided mergingReduces dependency on pre-specified cluster count k

Effective and General Distance Computation for Approximate Nearest Neighbor Search

Apr 25, 2024
MY
Mingyu Yang
🏛️ The Hong Kong University of Science and Technology | University of Leicester | Ant Group

To address the high computational cost, low accuracy, and poor generalizability of distance computation in high-dimensional approximate k-nearest neighbor (AKNN) search, this paper proposes a data-distribution-aware orthogonal projection distance estimation method with a decoupled, data-driven correction scheme. It is the first to incorporate explicit data distribution modeling into orthogonal projection-based distance estimation and fully decouples distance approximation from correction—thereby jointly optimizing efficiency, accuracy, and generality. The approach integrates orthogonal projection for dimensionality reduction, a lightweight data-driven correction model, high-dimensional index optimization, and accelerated distance computation mechanisms. Extensive experiments on multiple real-world datasets demonstrate that our method achieves 1.6–2.1× higher retrieval speed than ADSampling, while significantly improving recall and distance estimation accuracy.

Approximate Nearest Neighbor SearchEfficiency and Accuracy BalanceHigh-dimensional Space

Explaining the Success of Nearest Neighbor Methods in Prediction

May 31, 2018
GH
George H. Chen
🏛️ Carnegie Mellon University | Massachusetts Institute of Technology

Despite widespread empirical success, the theoretical foundations and practical deployment guidelines for k-nearest neighbors (k-NN) in predictive tasks remain inadequately understood. Method: We establish the first non-asymptotic error bound framework tailored for real-world deployment, replacing conventional smoothness or margin assumptions with verifiable cluster structure as the key success criterion. We integrate approximate nearest neighbor techniques—including LSH and graph-based indexing—and unify k-NN theory with emerging paradigms such as random forests, graphon modeling, and crowdsourcing. We further introduce a novel distance-learning perspective, characterizing how ensemble methods implicitly learn neighborhood structure. Contribution/Results: Evaluated on time-series forecasting, recommender systems, and medical image segmentation, our framework demonstrates that high accuracy is guaranteed solely under cluster-structured data—enhancing both theoretical interpretability and engineering tractability. It provides actionable, error-tolerance-driven guidance for data volume and hyperparameter selection.

Covers theoretical and practical applications.Explains success of nearest neighbor methods.Studies prediction in diverse real-world scenarios.

This study addresses the lack of theoretical guarantees for $k$-nearest neighbors (kNN) regression when applied to survey data arising from complex sampling designs, which violate the standard i.i.d. assumption. The paper establishes the first consistency framework for kNN regression under such settings by integrating probability sampling theory, nonparametric regression, and asymptotic analysis. It rigorously derives a lower bound on the convergence rate and demonstrates that kNN regression remains consistent even when accounting for sampling weights and dependence structures inherent in complex surveys. However, the convergence rate is shown to be adversely affected by the curse of dimensionality. These theoretical findings extend the classical kNN consistency results—previously limited to i.i.d. data—to realistic survey contexts. The validity of the theoretical conclusions is corroborated through both simulation studies and analyses of real-world survey data.

complex survey designsconsistencycurse of dimensionality

Post-clustering Inference under Dependency

Oct 18, 2023
JG
Javier González-Delgado
🏛️ Université de Rennes | ENSAI | CNRS | CREST-UMR 9194 | LAAS-CNRS | Université de Toulouse | Institut de Mathématiques de Toulouse | UMR5219 | UPS

Gao et al. (JASA 2022) proposed a post-clustering inference framework limited to i.i.d. Gaussian data, failing to accommodate arbitrary dependency structures among observations and features. Method: We generalize their framework to arbitrary dependence settings by developing a unified, dependence-aware post-clustering inference methodology compatible with hierarchical agglomerative clustering (single/complete/average linkage) and k-means. We derive theoretical conditions for well-defined p-values that ensure selective Type I error control and enable consistent covariance matrix estimation. Contribution/Results: Integrating selective inference, high-dimensional statistics, and covariance structure modeling, we design a robust testing pipeline. Experiments on synthetic and real-world protein structural data demonstrate substantial improvements in statistical reliability and practical utility for testing mean differences between clusters under dependence.

Establishing conditions for covariance matrix estimation compatibilityExtending post-clustering inference to dependent data structuresGeneralizing framework for hierarchical and k-means clustering algorithms

Latest Papers

What's happening recently
View more

This study addresses the common trade-off in post-processing ensemble forecasts, where neural network–based calibration often sacrifices predictive sharpness—particularly for short lead times—to improve reliability. To jointly optimize both calibration and sharpness, the authors propose a novel approach that introduces, for the first time, a sharpness penalty term directly into the continuous ranked probability score (CRPS) loss function. Assuming a Gaussian predictive distribution, the method is evaluated on ECMWF 2-meter temperature ensemble forecasts and achieves a reduction of 8.2%–12.5% in the width of central prediction intervals while maintaining CRPS and mean RMSE at levels comparable to baseline methods. This demonstrates a significant enhancement in forecast sharpness without compromising overall forecasting skill.

ensemble forecastneural networkpost-processing

This work investigates the interplay between interpolation and aggregation in regression tasks and its implications for learnability. By introducing the γ-graph dimension, the study characterizes the learnability boundary for a broad class of natural aggregation procedures and proposes a minimalist aggregation method that takes the median of three interpolating hypotheses. Theoretical analysis demonstrates that this median aggregation achieves optimal sample complexity among all finite interpolating aggregations and strictly outperforms standard interpolating learning. Moreover, the work reveals that certain hypothesis classes are learnable only via infinite or non-interpolating aggregations, thereby establishing fundamental limitations and optimality conditions for finite interpolating aggregation schemes.

aggregationinterpolationlearnability

This work addresses the problem of aggregating multiple calibrated Bayesian expert forecasts to construct a new predictor that remains calibrated and is Blackwell-dominated by the target expert, rather than merely minimizing a specific loss function. Under the setting where only the experts’ prior distributions are observed—without access to the true state—the authors formally define the aggregation objective as simultaneously achieving calibration and Blackwell refinability. By modeling calibrated experts through reduced-form information structures, they characterize the set of feasible predictions using the row space of a linear system intersected with a non-negative cone, and analyze it via Blackwell dominance theory. Their main contributions include efficient solvability of both randomized aggregation problems, while showing that determining the existence of a deterministic aggregator is NP-hard and admits no multiplicative PTAS unless P = NP, thereby revealing a fundamental computational distinction between randomized and deterministic aggregation.

Bayesian expertsBlackwell refinementcalibrated forecasts

This work proposes a general, model-agnostic post-processing framework that systematically integrates ensemble learning into fairness optimization. By aggregating predictions from multiple base models without requiring access to their internal architectures, the approach uniformly supports diverse predictive tasks—including classification, regression, and survival analysis—and accommodates a wide range of fairness definitions. The method addresses the inherent trade-off between predictive performance and fairness in machine learning models. Experimental results demonstrate that it significantly enhances fairness while maintaining or only marginally compromising prediction accuracy, thereby validating its effectiveness and generalizability across multiple scenarios.

fairnessmachine learningmodel fairness

Hot Scholars

HJ

Hyundong Jin

Ph.D candidate student in Yonsei Univ.
Language Model for CodeLanguage Model with RegexMachine LearningComputer Vision
YS

Yo-Sub Han

School of Computing, Yonsei University
automata theoryformal languagesalgorithm designinformation retrieval