subdifferential calculus

Designs and analyzes rules and formulas for computing subdifferentials and subgradients of nonsmooth functions and constructs the associated subdifferential sets. Uses these objects to derive optimality and stationarity conditions, relate primal and dual variables, and characterize the behavior of minimizing flows or other variational dynamics.

subdifferentialcalculus

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Some Primal-Dual Theory for Subgradient Methods for Strongly Convex Optimization

May 27, 2023
BG
Benjamin Grimmer
🏛️ Johns Hopkins University

This paper addresses strongly convex yet nonsmooth and non-Lipschitz optimization problems. We develop a unified primal–dual theoretical framework that, for the first time, reveals the equivalent dual-averaging representations of the subgradient method, proximal subgradient method, and switching subgradient method. Through a novel *dual-gap convergence analysis*, we establish the first $O(1/T)$ convergence guarantee applicable to this problem class. We derive an optimal stopping criterion and optimality certificate that require no additional computation, and rigorously characterize a controllable convergence boundary—even under early-stage exponential divergence. Our theory accommodates a broad range of step-size choices and accommodates ill-conditioned non-Lipschitz structures. While preserving algorithmic simplicity, our framework substantially extends both the applicability and theoretical depth of subgradient-type methods.

Convergence AnalysisConvex OptimizationSubgradient Methods

Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex functions

Mar 18, 2024
SZ
Siyuan Zhang
🏛️ Chinese Academy of Sciences | University of Chinese Academy of Sciences | National University of Singapore

This work studies the convergence of decentralized stochastic subgradient methods (DSGD-type algorithms) for nonsmooth, nonconvex objective functions that violate Clarke regularity—such as neural networks with non-differentiable activations (e.g., ReLU). We propose a unified analytical framework that, for the first time without assuming Clarke regularity, establishes asymptotic convergence guarantees for mainstream variants including DSGD, DSGD-T, and DSGD-M. Our analysis couples discrete iterations with differential inclusions and employs a coercive Lyapunov function to characterize the stable set; under mild regularity conditions and diminishing step sizes, we prove that iterates converge almost surely to this stable set. The framework accommodates both gradient tracking and momentum mechanisms. Numerical experiments on nonsmooth distributed neural network training confirm both theoretical reliability and practical efficiency of the proposed approach.

Convergence analysis without Clarke regularityDecentralized optimization for nonsmooth nonconvex functionsTraining nonsmooth neural networks efficiently

Escaping Saddle Points for Nonsmooth Weakly Convex Functions via Perturbed Proximal Algorithms

Feb 04, 2021
MH
Minhui Huang
🏛️ University of California, Davis

For nonsmooth weakly convex optimization, existing methods struggle to escape strict saddle points—hindering convergence to local minima. Method: This paper proposes a family of perturbed proximal algorithms—including perturbed proximal point, proximal gradient, and proximal linear variants—to address this challenge. Contributions/Results: We establish, for the first time, a verifiable characterization of ε-approximate local minima for nonsmooth weakly convex functions and provide the first theoretical guarantee for escaping strict saddle points. By integrating perturbed optimization, nonsmooth analysis, and saddle-point escape theory, all three algorithms achieve an iteration complexity of O(ε⁻² log d) under standard assumptions to compute an ε-approximate local minimum. This yields the first polynomial-time convergence guarantee for escaping saddle points in nonsmooth optimization.

Achieving ε-approximate local minimum efficiently in high dimensionsDeveloping perturbed proximal algorithms for nonsmooth optimizationEscaping saddle points for nonsmooth weakly convex functions

On the Complexity of Finding Small Subgradients in Nonsmooth Optimization

Sep 21, 2022
GK
Guy Kornowski
🏛️ Weizmann Institute of Science

This paper studies the oracle complexity of finding a (δ,ε)-stable point—a point whose δ-neighborhood contains a subgradient of norm at most ε—in nonsmooth optimization. Under Lipschitz continuity, it establishes for the first time that deterministic first-order algorithms necessarily incur dimension-dependent complexity, precluding dimension-free bounds; in contrast, randomized algorithms achieve the tight upper bound Õ(1/(δε³)), matched by a universal randomized lower bound. It further reveals that convexity dramatically accelerates convergence: for convex functions, a deterministic O(1/ε²) upper bound is attained, and it is proven that zero subgradients cannot be identified exactly in finite time; for smooth functions, derandomization is achievable with only logarithmic overhead. The core contribution lies in precisely characterizing the fundamental roles of determinism vs. randomness, convexity, and smoothness in governing oracle complexity, and providing tight upper and lower bounds for each setting.

Establish lower bounds for finding stationary points in any randomized algorithmImprove convergence rates for convex functions in finding stationary pointsStudy deterministic and randomized algorithms for finding stationary points in nonsmooth optimization

Geometry, Computation, and Optimality in Stochastic Optimization

Sep 23, 2019
CC
Chen Cheng
🏛️ Stanford University

This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.

Characterize optimality of stochastic gradient methods via geometryDetermine when nonlinear updates are necessary for optimal convergenceQuantify sub-optimality of subgradient methods using constraint convexity

Latest Papers

What's happening recently
View more

This work addresses the computation of strongly stationary solutions in stochastic convex optimization, where strong stationarity is defined by the presence of small elements near zero in the subdifferential. The paper introduces, for the first time, a notion of strong stationarity based on this criterion and employs tools from dimension theory to analyze the structure of subdifferential graphs, revealing how random sampling effectively preserves their essential features. Building on these insights, a proximal-point-type algorithm is developed, leveraging the Moreau envelope and subdifferential analysis to establish rigorous convergence guarantees. The proposed method theoretically overcomes the challenge posed by the lack of uniform convergence of subgradients in neighborhoods of optimal solutions, thereby enabling efficient approximation of strongly stationary points.

Moreau envelopeproximal-point methodsstationary point

This work addresses a class of nonconvex constrained optimization problems where both the objective and inequality constraints are compositions of convex Lipschitz outer functions with smooth inner mappings. The authors propose a smoothed proximal linear augmented Lagrangian method, reformulating the original problem as a nonsmooth nonconvex-concave minimax problem by restricting dual variables to a compact set. A finite-step mechanism is designed to map stationary points of the truncated minimax problem to KKT points of the original problem. Under a local cone regularity condition, they show that the artificial dual truncation automatically deactivates near feasible points, thereby establishing explicit convergence rates for the KKT residual: a global rate of $O(K^{-1/3})$ under dual regularization, which improves to $O(K^{-1/2})$ when the outer functions are piecewise linear and a local dual error bound holds.

augmented Lagrangiancomposite constraintsconstraint violation

This work addresses nonconvex, nonsmooth stochastic optimization problems where the objective comprises an expected loss and a lower semicontinuous regularizer, without assuming convexity or uniformly bounded gradient variance. The authors propose a proximal stochastic subgradient method incorporating an Armijo-type line search, which leverages progressively refined sample average approximations. Under the mild condition that the sample size is nondecreasing and diverges to infinity, they establish—within the nonsmooth Kurdyka–Łojasiewicz (KL) framework—the almost sure convergence of function values, global trajectory convergence, and the property that all cluster points are stationary. Notably, this is the first such result in the nonsmooth KL setting. Furthermore, for functions satisfying a KL inequality with exponential desingularizing functions, they derive a polynomial convergence rate up to a logarithmic factor, thereby relaxing conventional requirements on regularizer convexity and variance boundedness.

Kurdyka-Łojasiewicz conditionnonconvex optimizationnonsmooth optimization

This work addresses the lack of theoretical guarantees for Schedule-Free optimization methods in non-convex settings, particularly regarding convergence and saddle-point escape. By constructing a continuous-time limit that corresponds to a non-autonomous ordinary differential equation, the authors develop a Lyapunov-based analytical framework. Within this framework, they establish—for the first time—that the standard Schedule-Free gradient descent and its stochastic variant achieve the optimal worst-case convergence rate among first-order methods, without requiring algorithmic modifications or strong assumptions. Moreover, they rigorously prove that these methods avoid strict saddle points. Bridging non-convex optimization, ODE modeling, and non-autonomous dynamical systems theory, this study provides the first comprehensive theoretical foundation for both convergence and saddle-point escape in Schedule-Free methods.

convergence ratefirst-order methodsnonconvex optimization

Hot Scholars

MC

Michael C. Fu

University of Maryland
simulation optimizationstochastic gradient estimationqueueing
MB

Martin Burger

Deutsches Elektronen-Synchrotron DESY und Universität Hamburg
MathematicsImaging
GS

Gabriele Steidl

TU Berlin
Computational harmonic analysisoptimizationimage processingmachine learning
TR

Tim Roith

Postdoc, Deutsches Elektronen-Synchrotron DESY
Mathematics
FB

Francesco Bullo

Professor of Mechanical Engineering, UC Santa Barbara
Systems and ControlMulti-Agent SystemsRobotic NetworksPower Systems