Score
Designs, implements, and analyzes algorithms and procedures that locate minima or maxima of objective functions subject to constraints across continuous, discrete, convex, nonconvex, stochastic, and combinatorial problem settings. Develops and selects optimization methods and techniques—such as gradient-based, second-order, proximal, subgradient, randomized, and integer-programming approaches—evaluating their convergence, complexity, numerical stability, and integration with solver software.
This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.
This work addresses optimization problems with convex constraints whose intersection is difficult to project onto, covering both strongly convex smooth and general nonsmooth convex settings. The authors propose a novel algorithm that integrates stochastic feasibility methods with (sub)gradient descent, wherein each iteration randomly samples a subset of constraints and employs an adaptive Polyak stepsize that requires no prior knowledge of problem parameters, complemented by iterate averaging. Theoretical analysis establishes linear convergence under strong convexity and a worst-case rate of $O(1/\sqrt{T})$ for general convex objectives, while the infeasibility measure decays geometrically almost surely. Numerical experiments on QCQP and SVM tasks demonstrate superior computational efficiency over existing methods, and under specific sampling strategies, the algorithm achieves optimal convergence rates.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
Conventional design of convex optimization algorithms is often ad hoc and lacks systematic principles. Method: This paper proposes a novel algorithm construction paradigm grounded in RLC circuit modeling: (i) formulate a continuous-time circuit dynamical system whose trajectories converge to the optimizer; (ii) apply automated symbolic discretization coupled with Lyapunov stability analysis to rigorously guarantee global convergence of the resulting discrete-time iterative algorithm. Contribution/Results: This work establishes the first systematic mapping from circuit physics to optimization algorithm design, enabling provably convergent translation from continuous dynamics to discrete algorithms. It uniformly reconstructs classical methods—including gradient descent and Nesterov’s accelerated gradient—and synthesizes multiple new variants, including distributed algorithms. All derived algorithms come with formal convergence proofs, demonstrating the framework’s generality, mathematical rigor, and practical applicability.
This work addresses the convergence guarantees of stochastic line search optimization for over-parameterized models under interpolation conditions. We establish a necessary and sufficient condition on the search direction—applicable to a broad class of methods—that ensures finite termination and bounded backtracking steps, and rigorously prove linear convergence under the Polyak–Łojasiewicz (PL) assumption. The condition unifies major first-order strategies—including momentum, conjugate gradient, and adaptive preconditioning—providing a verifiable theoretical foundation for their principled integration with stochastic line search. Our analysis fills a critical gap in the convergence theory of stochastic line search methods and significantly extends both the applicability and reliability of efficient first-order optimization in interpolation learning regimes.
This work addresses the issue that quadratic penalty relaxations of binary linear programs often yield spurious or infeasible local minima. To overcome this, we propose a class of QUBO relaxation models satisfying specific structural conditions that guarantee all local minima are feasible and strictly binary. Leveraging these conditions, we derive novel differentiable relaxations for classical combinatorial optimization problems—including open-pit mining, the 0–1 knapsack problem, and the traveling salesman problem—and solve them using gradient-based optimizers such as projected gradient descent and Adam. Experimental results demonstrate that the proposed approach reliably converges to valid binary solutions, thereby establishing clear theoretical guarantees and delineating the applicability boundaries of differentiable optimization as a local solver for combinatorial problems.
This study investigates the statistical properties of Lagrange multipliers in constrained maximum likelihood estimation and least squares problems, along with their implications for numerical optimization. Leveraging large-sample theory, it establishes that under correctly specified models, Lagrange multipliers converge in probability to zero as the sample size grows, a result extended to high-dimensional settings such as deep learning. Building on this asymptotic behavior, the work provides the first statistical justification for initializing Lagrange multipliers at zero and integrates this insight into constrained optimization algorithms, including augmented Lagrangian methods and sequential quadratic programming. Numerical experiments demonstrate that this initialization strategy substantially enhances algorithmic stability and convergence efficiency in applications such as constrained regression and dynamic discrete choice models.
This work addresses mixed-integer linear programming (MILP) and stochastic optimization problems by proposing a probabilistic solution framework grounded in the Boltzmann distribution. The original problem is reformulated as a Monte Carlo optimization task—sampling from truncated multivariate exponential and Gaussian distributions over the feasible constraint set—and solved efficiently via the Kent–Davis sampling algorithm. Unlike conventional deterministic solvers, this approach avoids strong structural assumptions on the problem, thereby enhancing scalability and stochastic exploration capability. Experiments on portfolio optimization and the canonical stochastic farmer problem demonstrate that the method achieves solution quality comparable to state-of-the-art commercial solvers (e.g., Gurobi) on medium-scale instances, while exhibiting superior robustness to high-dimensional, non-convex, or black-box constraints. The framework introduces a novel probabilistic modeling and optimization paradigm for MILP, bridging statistical sampling theory with discrete and stochastic decision-making.
This work addresses the computational inefficiency in solving mixed-integer convex optimization problems involving binary indicator variables that govern continuous variables. To tackle this challenge, the authors propose the Coordinate Optimality Reconstruction (CORe) framework, which uniquely integrates coordinate-wise optimality conditions into the modeling of indicator variables. By combining closed-form characterizations with disjunctive reformulation techniques, CORe constructs a novel mixed-integer convex programming formulation that effectively exploits exploitable structures embedded in the problem’s sparsity pattern. The approach preserves global optimality while substantially enhancing the performance of branch-and-bound algorithms. Experimental results demonstrate that, across multiple problem classes—including quadratic programs and robust single-index models—CORe significantly accelerates solver convergence compared to conventional big-M formulations.