Score
Design, build, and analyze kernel functions and systems that measure similarity, weight or smooth values across inputs (e.g., feature vectors, tokens, graph nodes, or particle pairs); this includes selecting kernel families, specifying and tuning kernel hyperparameters, implementing kernel-based regression/weighting/smoothing, and optimizing kernel computation and memory. Also develop and evaluate specialized kernels and analyses—such as neural tangent kernels, particle-to-particle kernels, and kernel-graph residual smoothing—to propagate local similarity, enforce graph-based smoothness, and integrate kernel operations into model training and inference.
Standard convolutional operations cannot be directly applied to graph-structured data due to its irregular, non-Euclidean topology. Method: This paper proposes Graph Kernel-driven Learnable Structural Convolution (GK-Conv), a purely structural, end-to-end modeling framework operating directly on non-Euclidean graph domains. GK-Conv eliminates explicit graph embedding and instead constructs a parameterized, structural convolutional operator grounded in generic graph kernel functions—enabling plug-and-play integration of arbitrary graph kernels and generating CNN-style, interpretable structural masks. The model is fully differentiable and optimized via ablation-guided hyperparameter analysis. Contribution/Results: GK-Conv achieves state-of-the-art performance across multiple graph classification and regression benchmarks, empirically validating the central claim that strong generalization can be attained using topology alone—without node or edge features.
In multi-task learning, there is no standardized, test-set-free metric to uniformly assess the generalization capability of representations for downstream kernel ridge regression (KRR) tasks. Method: We propose the Uniform Kernel Prober (UKP), the first task-agnostic, kernel-driven pseudo-metric for representation quality. UKP explicitly encodes desired invariances via a user-specified kernel and estimates representation quality solely from training data. It enjoys an $O(1/sqrt{n})$ convergence rate and admits efficient computation. Contribution/Results: Unlike conventional metrics, UKP does not require task-specific labels or test sets, yielding comparable upper bounds on KRR generalization error across diverse features or representations. Extensive benchmark experiments demonstrate that UKP robustly discriminates representation quality and exhibits strong correlation with actual KRR generalization error—validating its effectiveness, robustness, and plug-and-play utility.
This work proposes a novel approach to graph classification by replacing the conventional parameterized classifier—such as a linear Softmax layer—in graph neural networks (GNNs) with non-negative kernel regression (NNK). Instead of relying on learnable parameters, the method constructs predictions via convex combinations of embeddings from similar training samples, effectively performing interpolation in the embedding space. This substitution not only enhances model interpretability by grounding predictions in actual training instances but also offers stronger theoretical guarantees for generalization. By eliminating the need for additional trainable parameters in the classification head, the approach provides a transparent and efficient mechanism that maintains predictive performance while improving the explainability and robustness of GNN-based graph classification.
This paper addresses the high computational cost and low statistical efficiency of kernel methods in supervised learning—particularly for tasks like Monte Carlo integration—by introducing the first adaptation of Kernel Thinning (KT) to supervised settings, yielding two novel estimators: KT-NW (Kernel Thinning–Nadaraya–Watson) and KT-KRR (Kernel Thinning–Kernel Ridge Regression). Methodologically, it designs regression-aware kernel functions to perform distribution-aware compression of labeled data, preserving essential statistical structure while drastically reducing dataset size. Theoretically, it establishes the first multiplicative error bound for KT tailored to supervised learning, jointly guaranteeing statistical accuracy and computational efficiency. Empirically, the proposed methods achieve quadratic speedups in both training and inference on synthetic and real-world benchmarks; they significantly outperform i.i.d. subsampling in statistical error and closely approximate the performance of full-data models.
While theoretical equivalence between the Laplace kernel and the Neural Tangent Kernel (NTK) is established only in the infinite-width limit, empirical validation under realistic finite-width and high-dimensional Euclidean space (ℝᵈ) settings remains lacking. Method: This work conducts the first systematic regression experiments to assess their practical equivalence, employing two rigorous criteria: exact kernel function matching and consistency of Gaussian process posterior predictions—both evaluated under finite-width and ℝᵈ conditions. Contribution/Results: We demonstrate strong empirical agreement between the Laplace kernel and NTK in regression performance and generalization behavior, even when idealized assumptions (e.g., infinite width) are relaxed. This robust similarity provides solid empirical justification for approximating the computationally expensive NTK with the efficient Laplace kernel. The findings advance the practical deployment of lightweight kernel methods in deep learning modeling, bridging theoretical insights with scalable real-world applications.
This work proposes a unified functional analytic framework that interprets both supervised and unsupervised learning as variational optimization problems within a function space induced by the data distribution. The central insight is that the fundamental distinction between these learning paradigms arises from the choice of the functional being optimized, rather than from differences in the underlying function space itself. Data structure is characterized via operators induced by the distribution, and target functions are estimated in the eigenbasis of these operators. This framework systematically integrates classical algorithms—including kernel methods, spectral clustering, and manifold learning—revealing their intrinsic coherence and underscoring the foundational role of function spaces and associated operators in modern machine learning.
This work addresses the challenge of predicting kernel regression learning curves solely from raw data statistics. We propose an analytical modeling framework grounded in the Hermite Eigenstructure Assumption (HEA), which enables exact inference of test risk as a function of sample size using only the empirical covariance matrix and a polynomial expansion of the target function. HEA provides, for the first time, analytic approximations of kernel eigenvalues and eigenfunctions for anisotropic real-world image data, uncovering a shared evolutionary mechanism between kernel learning and MLP feature learning in the Hermite polynomial basis. Integrating covariance analysis, target function decomposition, and kernel learning theory, our method achieves end-to-end learning curve prediction. Extensive validation on CIFAR-5m, SVHN, and ImageNet demonstrates high accuracy. Crucially, we empirically confirm that MLP feature learning strictly follows the hierarchical progression of Hermite polynomial orders predicted by HEA.
This work addresses the challenge of resource allocation among the number of training samples (N), input observation points (n), and output resolution (m) in operator learning. The authors propose a two-stage sampling framework: in the offline stage, a discrete representation of the operator is learned via kernel regression; in the online stage, the output function is reconstructed from predicted observations, enhanced by physics-informed constraints to improve accuracy. The study establishes a novel quantitative scaling law and error decomposition mechanism linking N, n, and m, and introduces a physics-informed online reconstruction strategy that avoids retraining. Theoretical analysis provides convergence guarantees and an error-balancing criterion, while numerical experiments validate the proposed scaling law and demonstrate the method’s superior performance in preserving physical consistency and achieving high reconstruction accuracy.
This work explores the effective integration of kernel methods into deep learning architectures by introducing “Sparse Kernels”—a differentiable, localized, and lazy variant of kernel ridge regression. The approach decouples feature representations, target values, and evaluation points into learnable or fixed parameters and implements them as modular, end-to-end trainable layers in PyTorch. By deferring training to inference time, enabling training-free transfer, and facilitating hybrid kernel–neural models, this method substantially expands the design space of deep learning. Experiments demonstrate that Sparse Kernel modules achieve performance comparable to neural readouts while significantly reducing training costs across convolutional networks, Vision Transformers, and reinforcement learning settings, and can also serve as plug-and-play components to enhance existing models.