kernel methods

Designing, tuning, and applying kernel functions (including bandwidth selection and aggregation) for nonparametric estimation and testing on heterogeneous data, and adapting kernel-based techniques to accelerate numerical solvers and statistical estimators.

kernelmethods

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

A Fast and Effective Large-Scale Two-Sample Test Based on Kernels

Oct 07, 2021
HS
Hoseung Song
🏛️ KAIST | University of California, Davis

To address the high computational cost, low statistical power, and bandwidth sensitivity of kernel two-sample tests on high-dimensional, large-scale data, this paper proposes a parameter-free robust kernel test. The method avoids bandwidth selection entirely while ensuring reliability and high power. Its core contributions are threefold: (1) a novel test statistic designed via theoretical analysis to eliminate power loss from data splitting; (2) a non-asymptotic significance control mechanism guaranteeing validity under finite samples; and (3) inherent suitability for high-dimensional settings, delivering uniformly high power across diverse alternative hypotheses. Experiments on synthetic and real-world datasets demonstrate that the proposed method achieves 10–100× speedup over MMD and state-of-the-art large-scale kernel tests, with average power gains of 15%–40%, all without any bandwidth tuning.

Addresses computational intensity of existing two-sample testsDevelops efficient kernel test for large-scale dataImproves power and robustness in high-dimensional settings

Application of Multivariate Selective Bandwidth Kernel Density Estimation for Data Correction

Jun 28, 2023
HB
H. Bui
🏛️ University of Bergen | Bergen Offshore Wind Centre

To address insufficient uncertainty quantification in data bias correction, this paper proposes Selective Bandwidth Kernel Density Estimation (SB-KDE), a multivariate KDE method. SB-KDE introduces a novel learnable selective KDE factor that jointly controls the scale and shape of multidimensional kernel functions. An adaptive bandwidth selection mechanism is incorporated, jointly optimizing bandwidth via two complementary criteria: Mean Conditional Squared Error (MCSE), which prioritizes correction accuracy by minimizing RMSE, and Least-Squares Cross-Validation (LSCV), which ensures overall probability density function (PDF) fidelity. Uncertainty in bias correction is explicitly modeled through the conditional PDF’s expectation and its associated confidence intervals. Extensive experiments on both synthetic and real-world datasets demonstrate that SB-KDE significantly outperforms conventional non-selective KDE methods, achieving simultaneous improvements in correction accuracy and uncertainty characterization.

Applies multivariate kernel density estimation for data correctionProposes selective bandwidth adjustment to improve accuracyQuantifies correction uncertainty using conditional PDF statistics

Kernel Ridge Regression Inference

Feb 13, 2023
RS
Rahul Singh
🏛️ Harvard University | Amazon Science

This paper addresses the challenge of statistical inference for kernel ridge regression (KRR) on non-standard data—such as preference rankings, graphs, and sequences—where conventional inferential frameworks fail. Method: We construct the first uniform confidence set for KRR with nearly minimax-optimal shrinkage rate. Our approach leverages a symmetric bootstrap procedure that automatically cancels bias while ensuring computational efficiency. Rigorous finite-sample coverage is established via reproducing kernel Hilbert space (RKHS) theory, uniform Gaussian/bootstrap coupling, and covering number analysis. Contribution/Results: The method enables direct hypothesis testing of matching effects—for instance, whether students benefit from attending higher-ranked schools—and is empirically validated in evaluating school assignment mechanisms. By providing valid, distribution-free inference for KRR on structured domains, our work fills a critical theoretical gap in econometrics and related fields, where KRR has seen widespread application but lacked formal inferential foundations.

Constructing uniform confidence bands for kernel ridge regressionEnabling inference with nonstandard data like preferences and graphsTesting match effects in school matching mechanisms

Approximation by non-symmetric networks for cross-domain learning

May 06, 2023
HM
H. Mhaskar
🏛️ Claremont Graduate University

Approximating high-dimensional, low-smoothness functions—common in cross-domain learning (e.g., invariant learning, transfer learning, SAR imaging)—remains challenging due to limitations of conventional symmetric or positive-definite kernels. Method: This paper proposes a neural network framework based on asymmetric, irregular kernels, breaking away from traditional kernel symmetry/positivity constraints. It systematically constructs generalized translation networks and rotationally banded function kernels—novel asymmetric kernel architectures—and establishes their approximation theory. It further introduces the ReLU<sup>r</sup> activation (with non-integer r > 0) and derives its uniform approximation error bound for Sobolev functions. Contribution/Results: Leveraging asymmetric kernel decomposition and Sobolev space analysis, the framework yields tight approximation error estimates for low-smooth, high-dimensional functions. Empirically, it achieves significantly improved cross-domain generalization accuracy under small-sample and low-regularity conditions.

High-Dimensional Smooth Function LearningIrregular Kernel NetworksReLU^r Activation Function

Practical Kernel Tests of Conditional Independence

Feb 20, 2024
RP
Roman Pogodin
🏛️ McGill University | Mila | Gatsby Computational Neuroscience Unit | University College London | University of British Columbia | Amii

Addressing the dual challenges of inflated Type I error rates (loss of test-level control) and low statistical power in conditional independence testing, this paper proposes a data-efficient kernel-based testing framework. The method employs kernel ridge regression and introduces, for the first time in this setting, three principled bias-correction strategies: data splitting, auxiliary data utilization, and restriction to simplified function classes—ensuring rigorous asymptotic and finite-sample control of the significance level. Theoretically, the approach guarantees convergence of the Type I error rate to the nominal significance level while enhancing detection power for complex dependency structures. Extensive experiments on diverse synthetic and real-world datasets demonstrate that the proposed method achieves precise Type I error control and substantially outperforms state-of-the-art competitors—including KCIT and RCIT—in statistical power, with improved robustness and reliability.

Addressing bias in test statistics from kernel ridge regression methodsDeveloping kernel-based tests for conditional independence with accurate false positive controlImproving test level accuracy while maintaining competitive statistical power

Latest Papers

What's happening recently
View more

This study addresses the limitation of the classical Weitzman overlap coefficient, which is restricted to pairwise probability distributions, by extending it for the first time to the case of k (k ≥ 2) independent distributions. The authors reformulate the generalized overlap coefficient as the expectation of a specific function and develop a nonparametric estimator that circumvents the need for closed-form density expressions by integrating kernel density estimation with the method of moments. The resulting framework offers a flexible and broadly applicable measure of overlap among multiple distributions. Extensive Monte Carlo simulations demonstrate that the proposed estimator exhibits robust performance across diverse distributional settings, combining strong theoretical validity with practical utility, thereby providing a valuable tool for multivariate overlap analysis.

distribution overlapk distributionskernel density estimation

This work addresses the challenge in kernel gradient descent algorithms where parameter selection relies heavily on cross-validation and lacks theoretical guarantees. To overcome this limitation, the authors propose an adaptive parameter selection strategy grounded in bias-variance decomposition and data splitting. By introducing an empirical effective dimension to quantify iterative increments, the method automatically adapts to diverse kernel functions, target functions, and error metrics. Within the framework of integral operators and statistical learning theory, the approach establishes optimal generalization error bounds. The proposed strategy is not only theoretically implementable but also achieves the minimax optimal convergence rate, offering a significant improvement over existing parameter selection methods in both theoretical rigor and practical performance.

adaptive strategygeneralization errorkernel-based gradient descent

This work addresses the challenge of resource allocation among the number of training samples (N), input observation points (n), and output resolution (m) in operator learning. The authors propose a two-stage sampling framework: in the offline stage, a discrete representation of the operator is learned via kernel regression; in the online stage, the output function is reconstructed from predicted observations, enhanced by physics-informed constraints to improve accuracy. The study establishes a novel quantitative scaling law and error decomposition mechanism linking N, n, and m, and introduces a physics-informed online reconstruction strategy that avoids retraining. Theoretical analysis provides convergence guarantees and an error-balancing criterion, while numerical experiments validate the proposed scaling law and demonstrate the method’s superior performance in preserving physical consistency and achieving high reconstruction accuracy.

budget allocationerror analysiskernel methods

Existing Bayesian optimization methods lack theoretical guarantees for adaptive data acquisition in nonlinearly parameterized models. This work proposes an analytical framework based on the reproducing kernel Hilbert space (RKHS) induced by kernels over the parameter space, integrated with regularized convex loss minimization, to establish a unified confidence bound theory for widely used nonlinear surrogate models. For the first time, this framework provides rigorous convergence guarantees for nonlinearly parameterized models under adaptive sampling, enabling a variety of novel acquisition strategies—including stochastic regularization and randomized model maximization—and substantially broadening the theoretical applicability of Bayesian optimization.

adaptive samplingBayesian optimizationkernel methods

This work addresses the lack of theoretical guarantees for nonparametric regression in reproducing kernel Hilbert spaces under model misspecification, high-dimensional settings, and nonconvex losses. It establishes a unified theoretical framework for regularized M-estimators encompassing a broad class of both convex and nonconvex loss functions. By introducing a novel complexity measure, the analysis achieves an explicit bias–variance decomposition. Leveraging tools from functional analysis and empirical process theory, the study proves the existence, measurability, and asymptotic linearity of the estimator without requiring closed-form solutions or global Lipschitz assumptions. Notably, within tensor-product Sobolev spaces, the framework reveals a mechanism to circumvent the curse of dimensionality, yielding minimax-optimal convergence rates that depend on mixed smoothness of the underlying function. The variance component is shown to be robust to model misspecification, and numerical experiments in C++ corroborate the theoretical findings.

convergence ratescurse of dimensionalityM-estimation

Hot Scholars

LB

Luca Benini

ETH Zürich, Università di Bologna
Integrated CircuitsComputer ArchitectureEmbedded SystemsVLSI
QL

Qian Lin

Research Engineer, ByteDance
DatabaseDistributed SystemData Streams
HO

Houman Owhadi

IBM Professor of Applied and Computational Mathematics and Control and Dynamical Systems. Caltech.
SciML. Kernel/GP Methods. UQ. Stochastic/Mulstiscale/Geometric Integration/Analysis..
AT

Akira Tamamori

Aichi Institute of Technology
Data MiningMachine LearningStatistics
AG

Arthur Gretton

Gatsby Computational Neuroscience Unit and Google Deepmind
generative modelscausalityhypothesis testingkernel methods