cyclic coordinate descent

Design and implement iterative optimization algorithms that minimize (including penalized) objective functions by updating parameters one coordinate at a time in a fixed cyclic order; build scalable solvers for high‑dimensional parameter estimation and analyze their convergence behavior, update rules, and computational complexity.

cycliccoordinatedescent

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.41
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes

May 28, 2024
JA
Jihao Andreas Lin
🏛️ University of Cambridge | MPI for Intelligent Systems

In large-scale Gaussian process (GP) hyperparameter optimization, iterative linear solvers—such as conjugate gradient (CG)—induce inefficiency in computing gradients of the marginal likelihood due to repeated, costly matrix-vector operations. Method: We propose a general-purpose optimization framework integrating pathwise gradient estimation, solver warm-starting, and budget-aware early stopping. The framework is agnostic to the underlying iterative solver and supports CG, alternating projections, and stochastic gradient descent. Contribution/Results: Our approach substantially alleviates the accuracy–efficiency trade-off in gradient estimation. Experiments demonstrate up to 72× speedup over standard CG when solving to full convergence. Under early stopping, the average residual norm drops to one-seventh of that achieved by baseline methods, significantly shortening hyperparameter optimization time while preserving convergence stability and gradient estimation accuracy.

Gaussian ProcessesHyperparameter OptimizationLarge-scale Datasets

Accelerated First-Order Optimization under Nonlinear Constraints

Feb 01, 2023
MM
Michael Muehlebach
🏛️ Max Planck Institute for Intelligent Systems | University of California

This paper addresses first-order optimization under nonlinear constraints—including nonconvex feasible sets—by proposing a novel accelerated algorithm grounded in nonsmooth dynamical systems. Methodologically, it models constraints in the **velocity space**, rather than the conventional position space, yielding sparse, local, and convex approximations of the feasible set and eliminating the need for expensive global projections at each iteration. Theoretically, the algorithm converges to stable points under nonconvex objectives and nonconvex constraints; under convexity, it achieves optimal acceleration rates in both continuous- and discrete-time settings. Its computational complexity scales nearly linearly with problem dimension and constraint count. Empirically, the method efficiently solves ℓ^p (p < 1) nonconvex regularized problems in compressed sensing and sparse regression: at p = 1, it matches state-of-the-art performance and substantially outperforms existing approaches.

Avoid full feasible set optimization per iterationDesign accelerated first-order algorithms for constrained optimizationHandle nonconvex constraints efficiently with sparse approximations

Iterative Linear Quadratic Optimization for Nonlinear Control: Differentiable Programming Algorithmic Templates

Jul 13, 2022
VR
Vincent Roulet
🏛️ Google Brain | University of Washington

This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.

Compare gradient descent, Gauss-Newton, Newton methodsOptimize nonlinear control using differentiable programmingTest algorithms on benchmarks like car racing

MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters

Feb 04, 2024
AS
Arsalan Sharifnassab
🏛️ University of Alberta | Leiden University

To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.

Dynamically adjusting step sizes during model optimizationOptimizing meta-parameters for efficient machine learning trainingReducing regret by considering long-term impact of learning rates

Geometry, Computation, and Optimality in Stochastic Optimization

Sep 23, 2019
CC
Chen Cheng
🏛️ Stanford University

This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.

Characterize optimality of stochastic gradient methods via geometryDetermine when nonlinear updates are necessary for optimal convergenceQuantify sub-optimality of subgradient methods using constraint convexity

Latest Papers

What's happening recently
View more

This work addresses the critical challenge that modern GPU-accelerated linear programming solvers—such as cuPDLP, which is based on the primal-dual hybrid gradient (PDHG) algorithm—exhibit performance highly sensitive to hyperparameters, yet lack tuning methods with provable generalization guarantees. For the first time, this study establishes structural relationships between hyperparameters and solution trajectories for multiple adaptive techniques in complex first-order LP solvers, including preconditioning, restart strategies, and smoothed weight updates. By integrating convergence analysis of PDHG with a model of structural sensitivity, the authors propose a data-driven hyperparameter learning framework that offers theoretical generalization guarantees under polynomial sample complexity. Experimental results demonstrate that the framework significantly enhances solver efficiency across diverse problem instances.

first-order methodsgeneralization guaranteesGPU acceleration

This work addresses the slow last-iterate convergence of the Extragradient method in unconstrained bilinear minimax optimization by introducing an acceleration mechanism based on dynamic stepsize scheduling. By formulating stepsize selection as an optimization problem, the paper establishes—for the first time—that convergence rates can be improved using only dynamically adjusted stepsizes, leading to a deterministic stepsize scheme following a power-law decay. The approach is further extended by allowing distinct stepsizes for the extrapolation and update steps, thereby approaching the optimal convergence rate. Under synchronized stepsizes, the method achieves a convergence rate of $O(T^{-2/3 + \varepsilon})$, which is shown to be tight in this setting; with asynchronous stepsizes, the rate improves to nearly optimal $O(T^{-1 + \varepsilon})$.

convergence ratedynamic stepsizesExtragradient method

This study addresses the computational challenges posed by high-dimensional regularized estimating equations, which often exhibit non-gradient structures, asymmetric Jacobians, over-identification, non-smoothness, non-convexity, or nested optimization, rendering standard penalized methods inefficient. To tackle this, the paper proposes a unified formulation of such problems as fixed-point equations and systematically develops four computational paradigms—minimization-based, Dantzig-type, regularization-based, and fixed-point-based—integrating strategies from penalized optimization, constrained linear programming, iterative root-finding, and proximal fixed-point iterations. This cohesive framework substantially enhances both solvability and algorithmic stability for high-dimensional regularized estimating equations, demonstrating broad applicability to complex settings such as longitudinal data analysis and survival modeling.

computational challengesestimating equationshigh-dimensional statistics

This study investigates the statistical properties of Lagrange multipliers in constrained maximum likelihood estimation and least squares problems, along with their implications for numerical optimization. Leveraging large-sample theory, it establishes that under correctly specified models, Lagrange multipliers converge in probability to zero as the sample size grows, a result extended to high-dimensional settings such as deep learning. Building on this asymptotic behavior, the work provides the first statistical justification for initializing Lagrange multipliers at zero and integrates this insight into constrained optimization algorithms, including augmented Lagrangian methods and sequential quadratic programming. Numerical experiments demonstrate that this initialization strategy substantially enhances algorithmic stability and convergence efficiency in applications such as constrained regression and dynamic discrete choice models.

asymptotic behaviorconstrained optimizationLagrange multipliers

Hot Scholars

CQ

Chenhao Qi

Southeast University
Signal Processing and Wireless Communications
MF

Moritz Flaschel

FAU Erlangen
computational mechanicsmaterial modelinginverse problemsmachine learning
TH

Trevor Hastie

Professor of Statistics, Stanford University
Statistical learning and modelingdata miningmachine learning
EK

Ellen Kuhl

Catherine Holman Johnson Director of Stanford Bio-X and Walter B. Reinhold Professor of Engineering
Automated ScienceMachine LearningAutomated Model DiscoveryLiving Matter