Score
Designs and implements methods to approximate positive-definite kernel functions and their associated reproducing kernel Hilbert spaces by constructing low-dimensional surrogate representations (e.g., Nyström low-rank decompositions, random feature maps) and choosing kernel parameterizations. Analyzes and balances approximation error, model complexity, and computational cost to produce efficient kernel-method implementations that preserve predictive behavior while reducing time and memory requirements.
Kernel ridge regression (KRR) suffers from prohibitive memory and computational costs in large-scale settings. This paper focuses on the low-rank approximation theory of KRR and makes four key contributions: (i) it derives, for the first time, a tight lower bound on the minimal rank required to preserve prediction consistency—providing rigorous, optimal theoretical guarantees for Nyström-type approximations; (ii) it proves that the computational complexity of the Nyström approximation is nearly linear in the number of samples; (iii) it establishes an approximation error bound for kernel functions in the range space of the integral operator and characterizes the growth behavior of the associated weight function norm; and (iv) it significantly expands the admissible range of regularization parameters. Collectively, these results unify the analytical framework for reliability, efficiency, and stability of low-rank KRR approximations, offering both foundational theoretical insights and practical guidance for scalable kernel learning.
Gaussian processes (GPs) struggle to rigorously incorporate uncountably infinite-dimensional functional prior information—such as boundary conditions or global physical constraints satisfied by PDE solutions. Method: This paper proposes a unified modeling framework grounded in reproducing kernel Hilbert spaces (RKHS), establishing for the first time a rigorous equivalence between the GP conditional expectation and orthogonal projection in RKHS. This enables direct embedding of functional constraints (e.g., Dirichlet or Neumann boundary conditions) into the GP prior, bypassing conventional pseudo-point approximations. Contribution/Results: We provide theoretical guarantees on existence, uniqueness, and convergence of the constrained GP posterior. Computationally, we design a practical numerical approximation algorithm. Experiments on PDE inverse problems demonstrate substantial improvements in uncertainty quantification accuracy and posterior consistency. The framework delivers a rigorous, general, and computationally tractable paradigm for integrating domain knowledge into Bayesian modeling.
Approximating high-dimensional, low-smoothness functions—common in cross-domain learning (e.g., invariant learning, transfer learning, SAR imaging)—remains challenging due to limitations of conventional symmetric or positive-definite kernels. Method: This paper proposes a neural network framework based on asymmetric, irregular kernels, breaking away from traditional kernel symmetry/positivity constraints. It systematically constructs generalized translation networks and rotationally banded function kernels—novel asymmetric kernel architectures—and establishes their approximation theory. It further introduces the ReLU<sup>r</sup> activation (with non-integer r > 0) and derives its uniform approximation error bound for Sobolev functions. Contribution/Results: Leveraging asymmetric kernel decomposition and Sobolev space analysis, the framework yields tight approximation error estimates for low-smooth, high-dimensional functions. Empirically, it achieves significantly improved cross-domain generalization accuracy under small-sample and low-regularity conditions.
This paper addresses the problem of efficiently approximating function integrals in a reproducing kernel Hilbert space (RKHS) given only i.i.d. samples from the target distribution. We propose a subsampling strategy based on (approximate) leverage scores—the first application of leverage scores to RKHS numerical integration—which drastically reduces the number of function evaluations required. Theoretically, we prove that only (m = O(log n)) subsampled points suffice to preserve the optimal (n^{-1/2}) convergence rate, and our error bound adapts to the smoothness of the integrand, achieving minimax-optimal rates in Sobolev spaces. Empirically, the method significantly improves the accuracy–efficiency trade-off over random or greedy quadrature on real-world datasets. Our results directly enable scalable computation of maximum mean discrepancy (MMD) and facilitate the design of efficient kernel-based hypothesis tests.
This work addresses the high computational complexity and lack of theoretical guarantees in multivariate (M ≥ 2) joint independence testing via the Hilbert–Schmidt Independence Criterion (HSIC). We propose the first Nyström-based HSIC estimator scalable to arbitrary M ≥ 2, built upon low-rank kernel matrix approximation and multivariate tensor kernel embeddings. We establish its statistical consistency under mild regularity conditions and reduce its time complexity from O(n²) to near-linear O(nm), where m ≪ n. Unlike existing methods—limited to bivariate (M = 2) settings and lacking theoretical foundations—our approach is the first to deliver a scalable, statistically consistent, and computationally efficient estimator for multivariate HSIC. Extensive experiments on synthetic data, media annotation dependency analysis, and causal discovery tasks demonstrate its effectiveness and practical utility.
This work addresses the challenge of incorporating prior features into kernel methods without penalizing them, thereby enhancing regression performance. To this end, the authors propose Conditional Kernel Ridge Regression (Conditional KRR), which decomposes the target function into a prior component modeled within a prescribed function class and a residual component, applying kernel regularization only to the latter. Theoretical analysis reveals that this approach is equivalent to standard Kernel Ridge Regression augmented with a controllable error term, and it achieves improved statistical risk under settings such as principal components or random features. When the prior component dominates the target function, Conditional KRR substantially outperforms standard KRR, a finding corroborated by both theoretical guarantees and empirical experiments.
This study addresses the high computational cost and energy consumption of supervised learning in the era of big data by investigating nonparametric supervised learning within a reproducing kernel Hilbert space. The authors propose a Horvitz–Thompson reweighted subsampling estimator based on empirical risk minimization. Through asymptotic analysis, they derive—for the first time—the optimal subsampling probabilities under the trace norm of the covariance operator and provide an efficient plug-in implementation. Both theoretical analysis and empirical evaluations demonstrate that the proposed method substantially reduces computational overhead while preserving estimation accuracy across synthetic and real-world datasets, offering an efficient and environmentally sustainable solution for large-scale nonparametric learning.
This work addresses the computational and memory bottlenecks of traditional Grassmannian kernel methods, which require constructing full Gram matrices and thus struggle with high-dimensional subspace data. To overcome these limitations, the authors propose a scalable kernel approximation framework based on random rank-one projections combined with bounded nonlinear transformations—either periodic or binary—that yield compact one-bit subspace feature representations. This approach enables continuous interpolation between the inverse Binet–Cauchy kernel and Gaussian-like kernels while effectively preserving the intrinsic geometry of subspaces. The method substantially reduces computational, memory, and storage costs. Experimental results on synthetic data and the ETH-80 classification benchmark demonstrate that the proposed technique accurately maintains Grassmannian geometric relationships with high fidelity, confirming its efficiency and practical utility.
This work addresses the scalability limitations of traditional kernel methods, which require constructing and inverting large kernel matrices, and the lack of generality in existing denoising approaches that often rely on restrictive assumptions about signals or noise. The authors propose an efficient operator learning algorithm based on Nyström subsampling for vector-valued regression in reproducing kernel Hilbert spaces, unifying denoising within a general operator learning framework. They innovatively relax classical Hölder-type and operator monotonicity constraints by introducing an indicator function to characterize more general source conditions, and for the first time apply Nyström approximation systematically to operator learning with functional outputs and universal denoising tasks. Experiments demonstrate that the method achieves performance comparable to full-kernel approaches at substantially reduced computational cost across diverse applications—including signal, audio, and image denoising, Radon inversion reconstruction, and energy efficiency prediction—while attaining minimax optimal convergence rates.
This study addresses the limitation of traditional Bayesian optimization to finite-dimensional vectors, which hinders black-box optimization in infinite-dimensional function spaces. To overcome this, we propose a functional Bayesian optimization method based on L0 manifold optimization. By jointly optimizing kernel locations and coefficients within sparse subspaces of reproducing kernel Hilbert spaces (RKHS), our approach effectively circumvents the dimensionality bottleneck. Key contributions include establishing a unified theoretical perspective that elucidates existing methods and developing a novel test function suite that extends finite-dimensional benchmarks to infinite-dimensional domains. Extensive experiments demonstrate that the proposed method achieves superior overall performance compared to state-of-the-art techniques across a wide range of benchmarks, thereby validating its effectiveness for optimization in infinite-dimensional spaces.