Score
Use the neural tangent kernel (NTK) framework to analyze and characterize the function approximation and training dynamics of wide or overparameterized neural networks. Concretely, derive approximation and error bounds, obtain score-approximation formulas, model SGD trajectories in the NTK regime, and quantify the effects of stochastic training noise and dataset reweighting on approximation and generalization.
This work addresses the challenge of extending Neural Tangent Kernel (NTK) theory—originally developed for regression—to classification settings, where cross-entropy loss typically drives logits to diverge, thereby violating the linearization assumption underpinning NTK. By introducing either parameter-space regularization or non-degenerate target conditions, the paper establishes, for the first time, sufficient conditions under which sufficiently wide neural networks maintain a "lazy training" regime in classification tasks, ensuring the NTK remains approximately constant throughout training. This advancement enables a rigorous extension of NTK theory to classification, allowing precise characterization of both training dynamics and generalization behavior. Moreover, it reveals a theoretical connection between the predictive distribution induced by random initialization and Bayesian inference.
This paper investigates whether the Neural Tangent Kernel (NTK) accurately characterizes the actual training dynamics of deep neural networks, particularly how its predictive error scales with network depth. Method: We conduct rigorous theoretical re-derivation, full-batch gradient descent simulations, trajectory tracking of generalization error across multilayer networks, and systematic comparisons against Gaussian process kernels. Contribution/Results: We find that NTK kernel regression exhibits significant deviation from the true training trajectories in both optimization and generalization behavior; remarkably, a simple Gaussian kernel achieves comparable performance, indicating that the NTK fails to deliver its theoretical advantages in practice. This work provides the first empirical evidence that the NTK’s theoretical equivalence to infinite-width networks breaks down substantially under standard training settings, challenging its ability to model path-dependent optimization effects. Our findings establish critical empirical bounds on the practical applicability of the NTK framework.
Conventional Neural Tangent Kernels (NTKs) require model parameters to be differentiable, limiting their applicability to nonsmooth objective functions, stochastic estimators, and nondifferentiable models. Method: We propose the Nonlocal Neural Tangent Kernel (NL-NTK), which replaces the classical gradient with a nonlocal gradient operator—thereby eliminating reliance on parameter differentiability. We construct two variants: a fixed-form NL-NTK and an attention-based dynamic NL-NTK, both grounded in nonlocal interaction operators that capture global dependencies in parameter space. Contribution/Results: Theoretical analysis and numerical experiments demonstrate that NL-NTK consistently characterizes training dynamics under gradient flow for both nonsmooth and stochastic models. It significantly extends the applicability of NTK theory beyond standard differentiable settings, offering enhanced theoretical generality and empirical robustness across diverse nonstandard learning scenarios.
This work unifies neural network learning and kernel learning theory by bridging the intrinsic connections between infinitely wide neural networks, the Neural Network Gaussian Process (NNGP), and the Neural Tangent Kernel (NTK). Method: We propose the Unified Neural Kernel (UNK), constructed as the inner product of gradient-descent-generated variables, jointly capturing training dynamics and initialization effects. UNK asymptotically unifies NNGP (the Bayesian zeroth-order limit) and NTK (the first-order tangent-space limit): it approximates NTK behavior in finite steps and converges to NNGP in the infinite-step limit. Theoretically, we establish uniform tightness and learning convergence guarantees for UNK, leveraging function-space analysis, random matrix theory, and gradient flow modeling. Results: Empirical evaluation across multiple benchmarks demonstrates that UNK significantly outperforms standalone NNGP or NTK, achieving superior generalization performance and enhanced training stability.
To address the failure of standard gradient-based training caused by non-differentiable activation functions (e.g., binary or spiking activations), this paper introduces the Surrogate Gradient Neural Tangent Kernel (SG-NTK), the first rigorous extension of NTK theory to training dynamics involving discontinuous activations and surrogate derivatives. Methodologically, leveraging functional analysis and kernel methods, we establish the existence and convergence of SG-NTK in the infinite-width limit and empirically validate its dynamical characterization capability on finite-width networks. Theoretically, we prove that SG-NTK exactly captures the surrogate gradient learning process—resolving the long-standing lack of theoretical foundation for surrogate gradient learning (SGL). Experimentally, on sign-activated networks, predictions derived from SG-NTK closely match those of kernel regression, confirming both theoretical soundness and practical applicability.
Existing Neural Tangent Kernel (NTK) theory relies on convergence bounds derived from the smallest eigenvalue, which are overly pessimistic and fail to explain the rapid convergence observed in practical neural network training. This work proposes a refined analytical framework based on the alignment between Label-NTK and Residual-NTK, revealing for the first time that the projections of labels and training residuals onto NTK eigenvectors scale proportionally with their corresponding eigenvalues. Leveraging this insight, the authors derive tight convergence and improved generalization bounds that depend on the full spectral structure of the NTK. Combining NTK linearized dynamics, spectral analysis, theoretical proofs, and extensive experiments across multiple datasets—including both MLPs and CNNs—the proposed bounds significantly outperform classical worst-case results, more accurately capture real-world training dynamics, and validate theoretical predictions on standard benchmarks.
This work addresses the challenge that the neural tangent kernel (NTK) of physics-informed neural networks (PINNs) often lacks guaranteed positive definiteness when solving linear partial differential equations, thereby hindering rigorous theoretical analysis of their training dynamics. To overcome this limitation, the authors propose a differential neural tangent kernel (DNTK) framework, which establishes, for the first time, a unified NTK analysis for PINNs incorporating a broad class of linear differential operators. They rigorously prove that the DNTK is positive definite in the infinite-width limit for both shallow and deep networks, under RePU and smooth non-polynomial activation functions. By integrating NTK theory, functional analysis, and differential operator theory, this study provides crucial theoretical foundations for the convergence of gradient-based optimization algorithms in PINNs.
This study addresses longstanding theoretical limitations in analyzing deep neural networks, specifically the reliance on overparameterization assumptions, high sample complexity, and the absence of parameter-level recovery guarantees. To overcome these challenges, this work proposes the ASPIRE algorithm for input-convex polynomial networks, which integrates active querying strategies, sampling-based diagonalization techniques, and iterative feature direction extraction to achieve precise layer-wise sampling. The primary contribution is the first exact parameter recovery guarantee under conditions where network depth grows exponentially with a polynomial. Furthermore, it is proven that all parameters can be recovered to δ-precision in polynomial time, substantially reducing sample complexity and establishing new theoretical bounds. These results rigorously validate the critical role of high-quality data in enabling efficient network training.
This work addresses the slow convergence of gradient descent on high-frequency targets in over-parameterized neural networks—a phenomenon attributed to spectral bias—and proposes a regularized Newton method to overcome this limitation. The authors introduce the “Neural Newton Tangent Kernel” (NNTK) to characterize the training dynamics of the proposed method in the infinite-width limit. By analyzing the spectral properties of the NNTK, they demonstrate that the regularization parameter uniformly controls the lower bound of its eigenvalues, thereby mitigating spectral bias. They further establish scaling rules for the regularization with respect to network width, ensuring a positive-definite Hessian and linearized training behavior. Theoretically, the method achieves global exponential convergence to a zero-loss solution for both low- and high-frequency targets in sufficiently wide networks, significantly outperforming standard gradient descent.
This work proposes the Distilled Neural Tangent Kernel (DNTK), a novel approach that integrates dataset distillation into the input space of the Neural Tangent Kernel (NTK) to address its high computational cost stemming from large Jacobian matrices. By combining Jacobian projection with low-rank approximation, DNTK substantially reduces computational complexity while preserving the kernel structure and predictive performance. Theoretical analysis and empirical results demonstrate that NTK matrices across various architectures exhibit low effective rank, which can be effectively retained through distillation. The method achieves up to five orders of magnitude reduction in NTK computation overhead and decreases Jacobian evaluation costs by 20–100×, striking a favorable balance between efficiency and fidelity.