Score
Designs and implements function approximators and estimators using reproducing kernel Hilbert spaces (RKHS), including constructing RKHS projections, analytic kernel estimators, and kernelized reward or value models. Analyzes their statistical and algorithmic properties — for example performing uncertainty quantification for exploration and deriving sample-complexity or regret guarantees using RKHS theory and reproducing kernel methods.
This paper addresses the long-standing conceptual divide between Gaussian processes (GPs) and reproducing kernel Hilbert space (RKHS) methods—two fundamental paradigms grounded in positive-definite kernels. We establish a unified theoretical framework rooted in the rigorous isometric isomorphism between the Gaussian Hilbert space and the RKHS. Our method formally proves equivalence between GP-based and RKHS-based solutions across diverse tasks: regression, interpolation, numerical integration, distributional discrepancy measurement (e.g., maximum mean discrepancy), statistical dependence quantification, and sample path analysis. Crucially, this framework bridges the epistemological gap between Bayesian probabilistic modeling and deterministic kernel methods. The results provide a coherent foundation for kernelized Bayesian inference, kernel manifold learning, and other cross-disciplinary applications—enabling principled integration of probabilistic and functional-analytic perspectives within kernel methods. (149 words)
This paper addresses the fragmentation and weak geometric intuition in existing kernel method theory by establishing a unified functional analytic framework grounded in Hilbert space geometry. Starting from the definition of positive-definite kernels, it rigorously unifies reproducing kernel Hilbert spaces (RKHS) and Hilbert–Schmidt operators, thereby reconstructing fundamental statistical concepts—including covariance, regression, and information-theoretic measures—within a coherent geometric setting. The work innovatively embeds kernel density estimation, distributional kernel embeddings, and maximum mean discrepancy (MMD) into a single RKHS paradigm, yielding a self-consistent theory bridging statistical estimation and probabilistic representation. This framework provides geometric interpretations for Gaussian processes and kernel Bayesian inference, and establishes a rigorous mathematical foundation for future theoretical advances in kernel-based learning.
This paper addresses the problem of efficiently approximating function integrals in a reproducing kernel Hilbert space (RKHS) given only i.i.d. samples from the target distribution. We propose a subsampling strategy based on (approximate) leverage scores—the first application of leverage scores to RKHS numerical integration—which drastically reduces the number of function evaluations required. Theoretically, we prove that only (m = O(log n)) subsampled points suffice to preserve the optimal (n^{-1/2}) convergence rate, and our error bound adapts to the smoothness of the integrand, achieving minimax-optimal rates in Sobolev spaces. Empirically, the method significantly improves the accuracy–efficiency trade-off over random or greedy quadrature on real-world datasets. Our results directly enable scalable computation of maximum mean discrepancy (MMD) and facilitate the design of efficient kernel-based hypothesis tests.
Gaussian processes (GPs) struggle to rigorously incorporate uncountably infinite-dimensional functional prior information—such as boundary conditions or global physical constraints satisfied by PDE solutions. Method: This paper proposes a unified modeling framework grounded in reproducing kernel Hilbert spaces (RKHS), establishing for the first time a rigorous equivalence between the GP conditional expectation and orthogonal projection in RKHS. This enables direct embedding of functional constraints (e.g., Dirichlet or Neumann boundary conditions) into the GP prior, bypassing conventional pseudo-point approximations. Contribution/Results: We provide theoretical guarantees on existence, uniqueness, and convergence of the constrained GP posterior. Computationally, we design a practical numerical approximation algorithm. Experiments on PDE inverse problems demonstrate substantial improvements in uncertainty quantification accuracy and posterior consistency. The framework delivers a rigorous, general, and computationally tractable paradigm for integrating domain knowledge into Bayesian modeling.
To address large integration errors and low posterior approximation accuracy in high-dimensional probability distribution compression, this paper proposes Generalized Kernel Thinning (GKT): a distribution compression framework operating in a reproducing kernel Hilbert space (RKHS) via flexible selection of target kernels, fractional-power kernels, or their weighted combinations. Our method establishes the first tight maximum mean discrepancy (MMD) error bound without requiring explicit square-root kernel representations; unifies treatment of both smooth and non-smooth kernels; and introduces the KT+ framework, jointly controlling function-level and distribution-level approximation errors. Experiments on 100-dimensional distributions and posterior compression for differential equation inference demonstrate significantly reduced integration error and MMD convergence rates that surpass those of Monte Carlo methods.
This study addresses the high computational cost and energy consumption of supervised learning in the era of big data by investigating nonparametric supervised learning within a reproducing kernel Hilbert space. The authors propose a Horvitz–Thompson reweighted subsampling estimator based on empirical risk minimization. Through asymptotic analysis, they derive—for the first time—the optimal subsampling probabilities under the trace norm of the covariance operator and provide an efficient plug-in implementation. Both theoretical analysis and empirical evaluations demonstrate that the proposed method substantially reduces computational overhead while preserving estimation accuracy across synthetic and real-world datasets, offering an efficient and environmentally sustainable solution for large-scale nonparametric learning.
Existing neural operators lack reliability and theoretical guarantees when handling out-of-distribution input functions. This work proposes an extended framework grounded in reproducing kernel Hilbert spaces (RKHS), leveraging kernel approximation techniques to achieve robust approximation of both out-of-distribution functions and their derivatives. The key innovation lies in establishing a theoretical connection between kernel selection and Sobolev eigenfunction spaces, thereby providing predictable guarantees on generalization error and derivative accuracy for neural operators. When applied to solving elliptic partial differential equations—particularly on manifolds represented as point clouds—the method demonstrates significantly enhanced geometric awareness, improved extrapolation accuracy, and greater computational efficiency.
Existing Bayesian optimization methods lack theoretical guarantees for adaptive data acquisition in nonlinearly parameterized models. This work proposes an analytical framework based on the reproducing kernel Hilbert space (RKHS) induced by kernels over the parameter space, integrated with regularized convex loss minimization, to establish a unified confidence bound theory for widely used nonlinear surrogate models. For the first time, this framework provides rigorous convergence guarantees for nonlinearly parameterized models under adaptive sampling, enabling a variety of novel acquisition strategies—including stochastic regularization and randomized model maximization—and substantially broadening the theoretical applicability of Bayesian optimization.
This work addresses the need for more efficient, robust, and flexible metrics for measuring distances between probability distributions in statistical inference and numerical integration. Centered on kernel methods, we propose an efficient estimator for Maximum Mean Discrepancy (MMD), develop novel MMD-based approaches for conditional expectation estimation and integral calibration, and introduce a new family of distance measures—kernel quantile discrepancies—that effectively overcome MMD’s limitations in tail sensitivity and discriminative power. Both theoretical analysis and empirical experiments demonstrate that the proposed methods offer strong scalability, computational efficiency, and superior performance, thereby providing more powerful and practical kernel-based tools for nonparametric statistics and integration tasks.
This work proposes a data-driven modeling approach for a broad class of nonlinear systems encompassing Volterra series, autoregressive models, and Hammerstein-type state-space realizations, without requiring explicit system identification. By extending Willems’ behavioral theory to vector-valued reproducing kernel Hilbert spaces (RKHS) and integrating minimal-norm interpolation with subspace identification techniques, the authors establish a unified framework for nonlinear system modeling. This study presents the first formulation of behavioral systems theory in vector-valued RKHS, thereby circumventing conventional identification procedures. The resulting framework enables direct application to simulation and control tasks and is applicable to a wide range of nonlinear dynamical systems.