Score
Designs and implements predictive models that combine learned neural-network feature maps with kernel methods—e.g., deep kernel learning (DKL) and neural tangent kernel (NTK) analyses—to produce regression and classification surrogates with principled uncertainty estimates. Builds and evaluates kernelized predictors, kernel approximations, and NTK-based linearizations to analyze generalization, identify epistemic uncertainty, and support uncertainty-aware ranking or decision workflows.
This paper investigates whether the Neural Tangent Kernel (NTK) accurately characterizes the actual training dynamics of deep neural networks, particularly how its predictive error scales with network depth. Method: We conduct rigorous theoretical re-derivation, full-batch gradient descent simulations, trajectory tracking of generalization error across multilayer networks, and systematic comparisons against Gaussian process kernels. Contribution/Results: We find that NTK kernel regression exhibits significant deviation from the true training trajectories in both optimization and generalization behavior; remarkably, a simple Gaussian kernel achieves comparable performance, indicating that the NTK fails to deliver its theoretical advantages in practice. This work provides the first empirical evidence that the NTK’s theoretical equivalence to infinite-width networks breaks down substantially under standard training settings, challenging its ability to model path-dependent optimization effects. Our findings establish critical empirical bounds on the practical applicability of the NTK framework.
Existing fixed-kernel theories—such as the Neural Tangent Kernel (NTK)—fail to capture the adaptive, dynamic feature learning underlying neural network generalization. Method: We propose an overparameterized Gaussian sequence model that admits closed-form characterization of feature evolution during training. Building upon this, we develop a statistical analysis framework that transcends the NTK regime. Contribution/Results: Our theory rigorously establishes quantitative links among feature evolution rate, representational capacity, and generalization error—without requiring infinite-width limits. It applies to practically sized architectures beyond the asymptotic “wide-network” regime, providing the first tractable theoretical prototype for adaptive feature learning in deep neural networks. By shifting focus from static kernel-based representations to dynamically evolving features, our work advances representation learning theory from the “fixed-kernel paradigm” toward a “dynamic-feature paradigm.”
This work challenges the paradigm of end-to-end training of deep neural networks (DNNs) by proposing **direct optimization of the Neural Tangent Kernel (NTK) to minimize generalization error**. Methodologically, it introduces a differentiable surrogate objective for generalization error—Kernel Alignment Risk Estimator (KARE)—and applies gradient descent to optimize NTK parameters explicitly, enabling kernel learning in the overparameterized regime. Theoretically, it integrates kernel alignment analysis to elucidate generalization mechanisms; empirically, it demonstrates on multiple synthetic and real-world benchmarks that KARE-optimized NTKs achieve stable generalization performance matching or significantly surpassing both the original DNN and the posterior (“after-kernel”) NTK. Crucially, this is the first work to elevate the NTK from a purely analytical tool to a trainable object, establishing a new paradigm for feature learning that bridges theoretical interpretability with strong empirical competitiveness.
This work addresses the challenge of extending Neural Tangent Kernel (NTK) theory—originally developed for regression—to classification settings, where cross-entropy loss typically drives logits to diverge, thereby violating the linearization assumption underpinning NTK. By introducing either parameter-space regularization or non-degenerate target conditions, the paper establishes, for the first time, sufficient conditions under which sufficiently wide neural networks maintain a "lazy training" regime in classification tasks, ensuring the NTK remains approximately constant throughout training. This advancement enables a rigorous extension of NTK theory to classification, allowing precise characterization of both training dynamics and generalization behavior. Moreover, it reveals a theoretical connection between the predictive distribution induced by random initialization and Bayesian inference.
This work unifies neural network learning and kernel learning theory by bridging the intrinsic connections between infinitely wide neural networks, the Neural Network Gaussian Process (NNGP), and the Neural Tangent Kernel (NTK). Method: We propose the Unified Neural Kernel (UNK), constructed as the inner product of gradient-descent-generated variables, jointly capturing training dynamics and initialization effects. UNK asymptotically unifies NNGP (the Bayesian zeroth-order limit) and NTK (the first-order tangent-space limit): it approximates NTK behavior in finite steps and converges to NNGP in the infinite-step limit. Theoretically, we establish uniform tightness and learning convergence guarantees for UNK, leveraging function-space analysis, random matrix theory, and gradient flow modeling. Results: Empirical evaluation across multiple benchmarks demonstrates that UNK significantly outperforms standalone NNGP or NTK, achieving superior generalization performance and enhanced training stability.
This work investigates the generalization performance of random feature methods under operator-valued kernels, with particular emphasis on the misspecified setting where the target function lies outside the associated reproducing kernel Hilbert space (RKHS). To this end, the authors develop a unified spectral regularization framework that encompasses both neural operators and neural networks within the neural tangent kernel (NTK) perspective for theoretical analysis. They extend random feature methods to operator-valued kernels for the first time and establish minimax optimal convergence rates in both well-specified and misspecified regimes. Key contributions include deriving optimal learning rates, quantifying the number of neurons required to achieve a prescribed accuracy, and strengthening the theoretical foundations of operator-valued kernel methods.
This work proposes the Distilled Neural Tangent Kernel (DNTK), a novel approach that integrates dataset distillation into the input space of the Neural Tangent Kernel (NTK) to address its high computational cost stemming from large Jacobian matrices. By combining Jacobian projection with low-rank approximation, DNTK substantially reduces computational complexity while preserving the kernel structure and predictive performance. Theoretical analysis and empirical results demonstrate that NTK matrices across various architectures exhibit low effective rank, which can be effectively retained through distillation. The method achieves up to five orders of magnitude reduction in NTK computation overhead and decreases Jacobian evaluation costs by 20–100×, striking a favorable balance between efficiency and fidelity.
This work addresses the critical limitation of existing deep learning weather forecasting models—their lack of reliable uncertainty quantification, which hinders high-stakes decision-making during extreme weather events. Leveraging the neural tangent kernel (NTK) framework, the authors introduce a Gaussian process correction term constructed from empirical features of the final network layer, enabling inference-time uncertainty estimation without model retraining. By uncovering an architecture-dependent variance collapse mechanism, they propose a data-driven decomposition strategy based on the spectral concentration of features, which for the first time yields prediction intervals that adapt to the severity of extreme events. At 90% coverage, the resulting intervals achieve 31–37% improved sharpness over split conformal prediction.
This work addresses the limitation of Bayesian Last Layer (BLL) methods, which perform Bayesian inference only on the final layer of a neural network and consequently underestimate uncertainty by neglecting epistemic uncertainty from preceding layers. To overcome this, the authors propose an improved approach that explicitly incorporates network-wide variability by projecting Neural Tangent Kernel (NTK) features into the last-layer feature space, provably enhancing posterior variance. A uniform subsampling strategy is introduced to mitigate computational overhead, accompanied by a theoretical bound on its approximation error. Empirical evaluations across UCI regression, contextual bandits, image classification, and out-of-distribution detection tasks demonstrate that the method significantly outperforms standard BLL and state-of-the-art baselines, achieving superior calibration and uncertainty estimation while maintaining computational efficiency.
This work addresses the challenge of balancing expressive power and computational efficiency in nonlinear models by proposing and open-sourcing “tnkm,” a JAX-based Python library that unifies nonlinear feature mappings with low-rank tensor network architectures. The framework offers the first scalable and modular implementation of tensor network kernel machines, enabling flexible composition of feature maps, network topologies, and optimization strategies—including alternating least squares and gradient-based methods. Experimental results demonstrate that the approach achieves competitive predictive accuracy on multiple nonlinear benchmark tasks while substantially reducing model parameter count and improving training efficiency. By providing a reproducible and high-performance platform, this work advances research in efficient nonlinear modeling.