Score
Design and apply analytical and computational methods that measure, predict, and explain the statistical properties (variance and higher moments) of gradients in parameterized models, with an emphasis on how those gradient statistics are localized across parameter neighborhoods; this includes constructing local two-copy/contraction analyses, accounting for parameter light-cone or neighborhood selection, and using local gradient moments to diagnose or predict gradient-related failure modes such as barren-plateau behavior.
This paper investigates the optimization and statistical behavior of Randomly Weighted Gradient Descent (RWGD) in linear regression. We analyze RWGD under general continuous weight distributions—including SGD and importance sampling as special cases—and establish the first non-asymptotic bounds on the convergence of both first- and second-order moments. We uncover an implicit regularization mechanism: the stationary distribution induced by gradient noise is characterized via geometric moment contraction, and we prove that weight design fundamentally governs the bias–variance trade-off. Crucially, we identify that rapidly convergent weight schemes inherently sacrifice statistical accuracy, and we provide the first quantitative characterization of the fundamental trade-off between convergence rate and estimation precision. Our results yield a unified analytical framework and concrete design principles for developing robust and efficient learning algorithms.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.
This work addresses the instability of parametric projections under input perturbations and the lack of effective evaluation metrics for local neighborhood stability. We propose a systematic assessment framework based on Gaussian perturbation probes, which— for the first time—integrates quantitative measures including displacement mean, variance, and nearest-anchor assignment error. Complementing these metrics, we introduce fine-grained visual diagnostics through displacement vectors, local PCA ellipsoids, and Voronoi misassignment maps. Evaluated on MNIST and Fashion-MNIST, our approach successfully identifies unstable regions that conventional reconstruction errors and neighborhood preservation metrics fail to detect, thereby establishing a new paradigm for robustness evaluation of projection methods.
Cross-validation (CV) for hyperparameter optimization in smoothing-penalized regression suffers from high computational cost and systematic underestimation of smoothing parameters due to unmodeled short-range autocorrelation. Method: We propose Neighborhood CV—a novel CV variant that combines QR decomposition and matrix calculus to compute CV gradients analytically and efficiently; it employs a neighborhood hold-out strategy with robust variance estimation, reducing hyperparameter optimization complexity to O(n), equivalent to a single model fit. Contribution/Results: This is the first method enabling efficient hyperparameter optimization and uncertainty quantification via neighborhood CV under short-range autocorrelation. Experiments demonstrate substantial improvements in accuracy and robustness of smoothing parameter estimation across smoothing quantile regression and GAMLSS models. Neighborhood CV establishes a scalable, theoretically grounded, and practical paradigm for hyperparameter selection in high-dimensional nonparametric modeling.
This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.
This study addresses the challenge of smooth modeling and geometric feature extraction for parameterized curves in $\mathbb{R}^p$ subject to discrete measurement errors. The proposed methodology leverages separable Hilbert spaces and Sobolev frameworks, employing penalized least squares to achieve smooth curve fitting. It further extends functional principal component analysis to $\mathbb{R}^3$, utilizing variational methods to solve for the eigenfunctions of the covariance operator in order to decompose spatial variance. By integrating the Euler–Lagrange theorem, regularization techniques, and differential geometry, this work effectively captures key differential features such as velocity and curvature. The resulting approach demonstrates significant advantages over conventional multivariate analysis methods.
This work investigates how training data shape the prediction mechanisms of neural networks through optimization trajectories, with a particular focus on higher-order effects under stochastic optimization and momentum. We introduce, for the first time, a second-order path kernel interpolation formula that expresses model predictions as an integral along the optimization path, where the leading term is weighted by the loss curvature and a correction term couples the covariance of gradient noise. This formulation naturally extends to momentum-based stochastic gradient descent. By leveraging path integrals, second-order Taylor expansions, and stochastic differential equation analysis, our framework precisely characterizes how stochasticity and momentum influence the interpolation structure and provides concentration bounds for the final prediction, quantifying the scale of predictive fluctuations.
In high-dimensional network analysis, when the number of observations is much smaller than the number of variables, model misspecification often leads to an inflated false positive rate in neighborhood estimation. This work proposes a conservative neighborhood selection method based on penalizing the volume of the model space, extending the Minimum Description Length (MDL) principle to high-dimensional settings and integrating ridge regression to elucidate its impact on mean squared error. The approach is applicable under both linear and nonlinear true models and theoretically guarantees either consistent neighborhood recovery under correct specification or a sparser—yet safer—estimate under misspecification, thereby substantially reducing false positive edges. This strategy overcomes key limitations of conventional methods such as Lasso and information criteria like AIC and BIC.
本文提出一种通过局部线性表示近似非线性机器学习模型来估计样本量的方法,以解决传统功效分析在复杂预测表面难以应用的问题。
This study addresses the absence of a unified theoretical framework and accuracy guarantees for gradient estimation via simulation output. To this end, it proposes the LTD (Learning to Differentiate) unified framework. Methodologically, by leveraging a differential interpretation of probability measure learning under weighted representations, this work translates model fitting accuracy into error bounds for gradient estimation. It further provides a unified convergence rate analysis for diverse methods, including kernel regression, local polynomial regression, kernel ridge regression, and smoothed neural networks. The primary contribution lies in achieving convergence rates comparable to standard Monte Carlo methods under smoothness conditions, thereby establishing a common theoretical foundation and rigorous accuracy guarantees for gradient estimation across multiple classes of learning approaches.