Score
Designs and analyzes algorithms and discretization schemes using smoothness measures averaged across coordinates or given per coordinate rather than a single global Lipschitz constant. This involves deriving average-coordinate smoothness bounds and coordinate-wise smoothness analyses to bound discretization error, enable variable step-size rules, and obtain nonasymptotic convergence or error-rate guarantees.
Classical iteration complexity analyses for first-order optimization methods rely on global Lipschitz continuity of the gradient, failing to exploit beneficial local smoothness—where the Lipschitz constant varies across regions—and thus incur unnecessary conservatism. Method: We introduce “glocal smoothness,” a novel structural assumption that simultaneously captures both global and local smoothness properties of the objective function—without dependence on algorithmic trajectories—thereby enabling trajectory-agnostic complexity bounds governed solely by intrinsic function constants. Contribution/Results: Under glocal smoothness, we establish improved iteration complexity for gradient descent with backtracking line search—surpassing that of fixed-step accelerated methods. Moreover, we provide a unified, refined convergence analysis for diverse algorithms including Polyak’s step size, adaptive gradient descent (AdGD), coordinate descent, stochastic and deterministic gradient methods, and nonlinear conjugate gradient, yielding significantly tighter complexity bounds across all cases.
This work addresses the slow convergence of gradient descent on complex objectives and its reliance on strong global smoothness assumptions. We introduce *directional smoothness*, a novel geometric concept characterizing the local smoothness of the objective function along the optimization trajectory—thereby circumventing restrictive global Lipschitz continuity requirements. Leveraging this path-dependent characterization, we derive a trajectory-aware suboptimality bound and formulate an implicit adaptive step-size equation. We theoretically establish that Polyak’s step size and normalized gradient descent inherently achieve path-adaptive fast convergence. Our methodology integrates directional smoothness analysis, implicit step-size design, and convergence theory for both convex and nonconvex settings. Experiments on logistic regression demonstrate that our new bound substantially improves upon classical $L$-smoothness-based guarantees. Notably, this is the first work to provide path-dependent convergence rates for these two canonical algorithms without requiring prior knowledge of smoothness parameters.
Existing theoretical frameworks struggle to characterize distribution-dependent generalization behavior of interpolation solutions in overparameterized models. Method: We introduce a distribution-dependent PAC-Chernoff bound—first enabling tight, precise generalization analysis of interpolation solutions—and define a computable model smoothness metric grounded in large-deviation theory. Building on this, we establish a unified theoretical framework linking regularization (ℓ₂, gradient penalty, initialization distance), data augmentation, and invariant architecture design to smoothness optimization. Results: We rigorously prove that prevalent training strategies—including weight decay, input gradient regularization, and data augmentation—implicitly enhance model smoothness. Our work provides the first distribution-dependent, tight, and interpretable theoretical foundation for the interpolation phenomenon in overparameterized learning, unifying empirical observations under a principled smoothness-centric lens.
We address convex optimization problems in machine learning that are nonsmooth yet satisfy $(L_0,L_1)$-smoothness—a structural generalization of classical $C^{1,1}$ smoothness. We propose a suite of novel algorithms that dispense with the standard $C^{1,1}$ assumption and achieve convergence rates independent of the initial point’s distance to the optimum. Methodologically, we (i) establish the first tight deterministic and stochastic convergence bounds for gradient clipping and the Polyak stepsize method, eliminating exponential dependence on initialization; (ii) design the first Nesterov-type accelerated algorithm for $(L_0,L_1)$-smooth convex functions, extended to stochastic overparameterized settings; and (iii) integrate Adaptive Gradient Descent (within the Malitsky–Mishchenko framework) for fully adaptive stepsize selection. All results hold for both strongly convex and general convex objectives, with rigorous theoretical guarantees. Our methods significantly improve upon state-of-the-art performance under nonsmooth yet structured smoothness assumptions.
Discrete approximations of learnable-scale Gaussian derivative filters in deep learning suffer from discretization inaccuracies at extremely small scales, leading to deviations from continuous-scale space theory. Method: This paper systematically investigates two hybrid discretization schemes—normalized-sampling Gaussian kernel convolution combined with central differencing, and integral Gaussian kernel convolution with central differencing. Contribution/Results: We quantitatively characterize, for the first time, their spatial smoothing magnitude and scale estimation consistency bias at infinitesimal scales, revealing the fundamental mechanisms behind their departure from continuous-scale space theory. Our work fills a critical methodological gap where discrete Bessel kernels fail. Furthermore, we establish a unified quantitative evaluation framework for multi-order spatial derivatives under consistent scale computation, rigorously distinguishing the intrinsic trade-offs between smoothing fidelity and scale consistency in both methods. This provides a theoretically sound yet practically viable discrete scheme for differentiable scale-space modeling.
This work addresses the looseness of existing non-asymptotic error bounds for Langevin Monte Carlo methods under strongly log-concave distributions, which overly rely on global smoothness constants and consequently deteriorate in high-dimensional or correlated-covariate settings. To overcome this limitation, the paper introduces a coordinate-wise averaged smoothness condition to characterize the potential function and combines synchronous coupling with Wasserstein distance analysis to derive substantially tighter bounds. The key innovation lies in replacing the global smoothness constant with an average coordinate-wise counterpart and employing a trace-type third-order smoothness quantity to weaken the Hessian-Lipschitz assumption. These improvements are extended to variable step sizes, Laplacian-smooth potentials, and finite-sum structures such as SGLD. Notably, the resulting bounds exhibit significantly improved dimension dependence in high-dimensional generalized linear models, especially under covariate correlation, offering broader theoretical applicability and outperforming current state-of-the-art results.
This work investigates the optimal query complexity lower bounds for first-order gradient oracles in finding ε-stationary points of higher-order smooth nonconvex optimization problems. By introducing a novel “block-chain” mechanism to construct hard instances and integrating higher-order smoothness analysis with deterministic first-order complexity theory, the study establishes the first dimension-free tight lower bounds: Ω(ε⁻⁷/⁴) under Hessian-Lipschitz continuity and Ω(ε⁻⁵/³) in the third-order smooth setting. These bounds match the known upper bounds achieved by existing accelerated algorithms, thereby resolving a long-standing gap in the theoretical understanding of nonconvex optimization complexity.
This work investigates the regularity control of vector fields and score functions in flow matching and diffusion models, aiming to derive sampling error bounds that do not deteriorate exponentially with dimension or spatial scale. Under general assumptions on the target distribution, the authors establish time- and dimension-optimal Lipschitz regularity estimates and introduce a one-sided Lipschitz condition to construct a globally Lipschitz transport map, thereby obtaining the first Wasserstein error bounds free from exponential degradation. By combining Euler discretization with functional inequalities—specifically Poincaré and log-Sobolev inequalities—they achieve an optimal sampling error rate of \(O(\sqrt{d}/N)\) (up to logarithmic factors) and establish corresponding functional inequalities for a broad class of probability measures.
This work investigates the generalization error and stability of gradient descent (GD) and stochastic gradient descent (SGD) under deterministic or random rounding in discrete parameter spaces. Leveraging frameworks of uniform stability and parameter uniform stability, and assuming convexity, Lipschitz continuity, and smoothness, the authors derive generalization bounds for both algorithms under rounding operations. Their main contributions include showing that deterministic rounding degrades GD’s generalization error to $O(T/\sqrt{n})$ and undermines its stability; demonstrating that SGD retains nontrivial stability under deterministic rounding, with bounds of $O(T/n)$ in one dimension and $O(T^2/n)$ in higher dimensions; and establishing a tight upper bound on parameter stability for random rounding under coordinate-wise separable losses.
This work addresses the dominant role of discretization error in reverse-time sampling of diffusion models under a fixed inference budget, a factor overlooked by existing non-asymptotic analyses that are often loose and ignore data structure. By deriving first-order asymptotic expansions for both the weak error of the Euler–Maruyama scheme and the Fréchet discretization error, the paper establishes—for the first time—an explicit connection between this error and intrinsic data geometry, such as the covariance spectrum, as well as the diffusion schedule. Under the exact score assumption, the integration of numerical analysis for stochastic differential equations with asymptotic theory yields a computable, geometry-aware optimization objective in the Gaussian setting. This formulation demonstrates strong predictive performance across diverse image generation and posterior sampling tasks, offering both theoretical grounding and practical tools for geometry-informed schedule design.