Second-Order Path Kernel Interpolation Formulas in Machine Learning

📅 2026-06-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how training data shape the prediction mechanisms of neural networks through optimization trajectories, with a particular focus on higher-order effects under stochastic optimization and momentum. We introduce, for the first time, a second-order path kernel interpolation formula that expresses model predictions as an integral along the optimization path, where the leading term is weighted by the loss curvature and a correction term couples the covariance of gradient noise. This formulation naturally extends to momentum-based stochastic gradient descent. By leveraging path integrals, second-order Taylor expansions, and stochastic differential equation analysis, our framework precisely characterizes how stochasticity and momentum influence the interpolation structure and provides concentration bounds for the final prediction, quantifying the scale of predictive fluctuations.
📝 Abstract
Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. In this paper, we develop second-order forms of these interpolation formulas. We show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling the curvature of the prediction with the covariance of mini-batch gradient noise. We also extend the representation to stochastic gradient descent with momentum, where the interpolation structure is preserved but with the weights modified by a memory-related factor. Moreover, we establish a concentration estimate for the terminal prediction, identifying the fluctuation scale around the expected second-order representation. Together, these results provide a refinement of the path-kernel interpretation of neural network prediction.
Problem

Research questions and friction points this paper is trying to address.

path kernel
second-order interpolation
stochastic gradient descent
neural network prediction
optimization path
Innovation

Methods, ideas, or system contributions that make the work stand out.

second-order interpolation
path kernel
stochastic gradient descent
curvature-weighted term
gradient noise covariance
🔎 Similar Papers
No similar papers found.