learning-enabled admm

Designs and analyzes ADMM-based optimization algorithms that incorporate learned modules (e.g., learned Moreau envelopes or MEL/sMEL variants) to reduce per-iteration computational cost and accelerate convergence while accommodating smooth and non‑smooth objective terms. Builds the algorithmic components and accompanying proofs or diagnostics to ensure the modified ADMM preserves convergence and feasibility guarantees for the targeted convex problem classes.

learning-enabledadmm

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses convex optimization problems involving both smooth and nonsmooth components by proposing the LEAF framework, which introduces scalar-valued Moreau envelope learning into the Alternating Direction Method of Multipliers (ADMM) for the first time. The approach explicitly models the Moreau envelope of the objective function using input-convex neural networks (ICNNs), yielding the MEL-ADMM algorithm and its splitting variant, sMEL-ADMM. These methods substantially reduce model complexity while rigorously preserving convexity and theoretical convergence guarantees. Experimental results demonstrate that the proposed algorithms achieve up to an order-of-magnitude speedup over existing solvers while maintaining a low optimality gap, and exhibit convergence rates comparable to those of classical ADMM.

accelerated optimizationADMMconvex optimization

This work proposes an online learning strategy for accelerating the convergence of the Alternating Direction Method of Multipliers (ADMM) when solving structured convex optimization problems—such as time-varying quadratic programs arising in model predictive control—by dynamically tuning its relaxation parameter. The approach targets scenarios where the problem structure remains fixed but parameters change over time, thereby circumventing the need for costly matrix refactorizations typically required in conventional penalty parameter adjustment. For the first time, convergence guarantees are established for ADMM with time-varying penalty and relaxation parameters. By integrating ideas from reinforcement learning into parameter scheduling, the method achieves substantial improvements in solution efficiency while maintaining low computational overhead. Implemented within the OSQP framework, the proposed strategy significantly reduces both iteration counts and actual solve times on standard quadratic programming benchmarks.

ADMMconvergence guaranteesonline learning

This work addresses the computational challenge in traditional alternating direction method of multipliers (ADMM) for bilinear minimax (saddle-point) optimization problems, where evaluating complex proximal operators is often required. The authors propose a novel ADMM variant that decomposes the original problem into two alternating substeps: a generalized projection onto the constraint set \( S \) and a Euclidean projection onto the set \( C \). The key innovation lies in the exact reformulation—without approximation or linearization—of the ADMM proximal operator under the bilinear structure into a computable generalized projection. By integrating tools from convex analysis and projection techniques, the method establishes a provably convergent and computationally efficient framework, significantly simplifying the solution process for bilinear minimax problems.

bilinear objectivesconvex optimizationminimax problems

We address convex optimization problems in machine learning that are nonsmooth yet satisfy $(L_0,L_1)$-smoothness—a structural generalization of classical $C^{1,1}$ smoothness. We propose a suite of novel algorithms that dispense with the standard $C^{1,1}$ assumption and achieve convergence rates independent of the initial point’s distance to the optimum. Methodologically, we (i) establish the first tight deterministic and stochastic convergence bounds for gradient clipping and the Polyak stepsize method, eliminating exponential dependence on initialization; (ii) design the first Nesterov-type accelerated algorithm for $(L_0,L_1)$-smooth convex functions, extended to stochastic overparameterized settings; and (iii) integrate Adaptive Gradient Descent (within the Malitsky–Mishchenko framework) for fully adaptive stepsize selection. All results hold for both strongly convex and general convex objectives, with rigorous theoretical guarantees. Our methods significantly improve upon state-of-the-art performance under nonsmooth yet structured smoothness assumptions.

Convergence RateMachine Learning OptimizationNon-smooth Functions

Directional Smoothness and Gradient Methods: Convergence and Adaptivity

Mar 06, 2024
AM
Aaron Mishkin
🏛️ Stanford University | Princeton University | Meta AI | Flatiron Institute

This work addresses the slow convergence of gradient descent on complex objectives and its reliance on strong global smoothness assumptions. We introduce *directional smoothness*, a novel geometric concept characterizing the local smoothness of the objective function along the optimization trajectory—thereby circumventing restrictive global Lipschitz continuity requirements. Leveraging this path-dependent characterization, we derive a trajectory-aware suboptimality bound and formulate an implicit adaptive step-size equation. We theoretically establish that Polyak’s step size and normalized gradient descent inherently achieve path-adaptive fast convergence. Our methodology integrates directional smoothness analysis, implicit step-size design, and convergence theory for both convex and nonconvex settings. Experiments on logistic regression demonstrate that our new bound substantially improves upon classical $L$-smoothness-based guarantees. Notably, this is the first work to provide path-dependent convergence rates for these two canonical algorithms without requiring prior knowledge of smoothness parameters.

Complex FunctionGradient DescentOptimization Efficiency

Latest Papers

What's happening recently
View more

Existing end-to-end approaches to solving constrained convex optimization problems often fail to strictly satisfy constraints and lack guarantees of optimality. This work proposes a trainable architecture based on unfolded ADMM that enforces hard constraints through an embedded constraint-satisfaction module and a differentiable equality-constraint correction layer, ensuring exact feasibility at every iteration. Furthermore, first-order optimality conditions are incorporated as soft constraints into the training objective to guide convergence toward high-quality solutions. The proposed method uniquely unifies strict constraint satisfaction with optimality-aware learning within an unfolded optimization framework. Empirical results across multiple constrained convex optimization tasks demonstrate substantial improvements over conventional black-box end-to-end models, achieving both high solution accuracy and strong constraint compliance.

black-box mappingconstrained convex optimizationconstraint satisfaction

This work addresses the limitations of traditional alternating direction method of multipliers (ADMM) and block coordinate descent (BCD) methods in large-scale optimization, particularly regarding computational efficiency and convergence guarantees. The authors propose an adaptive proximal ADMM algorithm along with two BCD variants that accommodate inexact subproblem solutions. A key innovation is the introduction of an inexact proximal mapping with dynamic error control, and the paper establishes, for the first time, convergence rate guarantees for stochastic BCD under Hölder smoothness assumptions. Theoretical analysis demonstrates that the proposed algorithms achieve optimal iteration complexity, matching the best-known rates for Lipschitz-smooth settings across nonconvex, convex, and strongly convex cases. Numerical experiments further confirm that the dynamic error strategy significantly outperforms fixed-error approaches.

ADMMblock coordinate descentblock decomposable methods

Alternating Direction Method of Multipliers for Nonlinear Matrix Decompositions

Dec 19, 2025
AA
Atharva Awari
🏛️ University of Mons

This paper addresses the nonlinear matrix decomposition (NMD) problem: given an input matrix $X$ and target rank $r$, find low-rank factors $W$ and $H$ such that $X approx f(WH)$, where $f$ is an element-wise nonlinearity. We systematically introduce the alternating direction method of multipliers (ADMM) into the NMD framework—enabling support for arbitrary differentiable or nondifferentiable nonlinearities (e.g., ReLU, square, MinMax) and composite loss functions (e.g., least squares, $ell_1$, KL divergence). Theoretically, we establish convergence guarantees for the proposed algorithm under nonconvex settings. Empirically, our method achieves significant improvements in accuracy and generalization across diverse tasks—including sparse nonnegative data fitting, probabilistic circuit modeling, and recommendation systems—while maintaining computational scalability and numerical stability.

Proposes ADMM algorithm for nonlinear matrix decompositionSolves approximation X ≈ f(WH) with element-wise nonlinear functionSupports diverse loss functions for real-world applications

This work addresses optimization problems with convex constraints whose intersection is difficult to project onto, covering both strongly convex smooth and general nonsmooth convex settings. The authors propose a novel algorithm that integrates stochastic feasibility methods with (sub)gradient descent, wherein each iteration randomly samples a subset of constraints and employs an adaptive Polyak stepsize that requires no prior knowledge of problem parameters, complemented by iterate averaging. Theoretical analysis establishes linear convergence under strong convexity and a worst-case rate of $O(1/\sqrt{T})$ for general convex objectives, while the infeasibility measure decays geometrically almost surely. Numerical experiments on QCQP and SVM tasks demonstrate superior computational efficiency over existing methods, and under specific sampling strategies, the algorithm achieves optimal convergence rates.

adaptive step sizesconstrained optimizationconvex constraints

This work addresses the limitations of the classical Frank-Wolfe method, which relies on a global linear minimization oracle (LMO) and requires bounded feasible sets along with curvature assumptions. The authors propose a Local LMO algorithm that solves a local linear minimization problem over the intersection of a neighborhood around the current iterate and the constraint set, thereby enabling projection-free gradient-type optimization. This approach extends the convergence theory of projected gradient descent to projection-free settings, eliminating the need for boundedness and curvature conditions. It provides a unified framework for convex, strongly convex, nonconvex, and stochastic optimization problems. Under various settings, the algorithm achieves optimal convergence rates matching those of projected gradient descent: sublinear for convex objectives, linear for smooth strongly convex cases, and optimal sublinear rates for nonconvex, stochastic, and nonsmooth scenarios.

constrained optimizationconvergence ratesgradient methods

Hot Scholars

SS

Shaojie Shen

Associate Professor, Hong Kong University of Science and Technology
Robotics
MJ

Mathews Jacob

University of Virginia
Image reconstructionMRIImage Analysis
JR

Jyothi Rikhab Chand

Department of Electrical and Computer Engineering, University of Virginia
Signal ProcessingOptimizationDeep Learning