Conditional KRR: Injecting Unpenalized Features into Kernel Methods with Applications to Kernel Thresholding

📅 2026-05-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of incorporating prior features into kernel methods without penalizing them, thereby enhancing regression performance. To this end, the authors propose Conditional Kernel Ridge Regression (Conditional KRR), which decomposes the target function into a prior component modeled within a prescribed function class and a residual component, applying kernel regularization only to the latter. Theoretical analysis reveals that this approach is equivalent to standard Kernel Ridge Regression augmented with a controllable error term, and it achieves improved statistical risk under settings such as principal components or random features. When the prior component dominates the target function, Conditional KRR substantially outperforms standard KRR, a finding corroborated by both theoretical guarantees and empirical experiments.
📝 Abstract
Conditionally positive definite (CPD) kernels are defined with respect to a function class $\mathcal{F}$. It is well known that such a kernel $K$ is associated with its native space (defined analogously to an RKHS), which in turn gives rise to a learning method -- called conditional kernel ridge regression (conditional KRR) due to its analogy with KRR -- where the estimated regression function is penalized by the square of its native space norm. This method is of interest because it can be viewed as classical linear regression, with features specified by $\mathcal{F}$, followed by the application of standard KRR to the residual (unexplained) component of the target variable. Methods of this type have recently attracted increasing attention. We study the statistical properties of this method by reducing its behavior to that of KRR with another fixed kernel, called the residual kernel. Our main theoretical result shows that such a reduction is indeed possible, at the cost of an additional term in the expected test risk, bounded by $\mathcal{O}(1/\sqrt{N})$, where $N$ is the sample size and the hidden constant depends on the class $\mathcal{F}$ and the input distribution. This reduction enables us to analyze conditional KRR in the case where $K$ is positive definite and $\mathcal{F}$ is given by the first $k$ principal eigenfunctions in the Mercer decomposition of $K$. We also consider the setting where $\mathcal{F}$ consists of $k$ random features from a random feature representation of $K$. It turns out that these two settings are closely related. Both our theoretical analysis and experiments confirm that conditional KRR outperforms standard KRR in these cases whenever the $\mathcal{F}$-component of the regression function is more pronounced than the residual part.
Problem

Research questions and friction points this paper is trying to address.

conditional kernel ridge regression
unpenalized features
kernel methods
residual kernel
Mercer decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

conditional kernel ridge regression
residual kernel
Mercer decomposition
random features
unpenalized features
💼 Related Jobs
No related jobs found.
R
Rustem Takhanov
1Department of Mathematics, Nazarbayev University, Astana, Kazakhstan 2Nazarbayev University Research Administration, Astana, Kazakhstan
Zhenisbek Assylbekov
Zhenisbek Assylbekov
Purdue University Fort Wayne
StatisticsNLPMachine Learning