Score
Design and analyze local-search optimization routines that use linear programming (LP) formulations and lifting techniques to compute local descent directions and candidate updates. Implement LP-based procedures that minimize scaled directional derivatives subject to training or feasibility constraints, preserve specified optimality conditions, and enable focused updates to selected model components.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This paper addresses parametric multiple-instance linear programming (LP), where the constraint matrix varies linearly across $p$ parameter values and the problem has $m$ constraints. We propose three matrix-decomposition-based warm-start algorithms—novel applications of eigenvalue decomposition, Schur decomposition, and an enhanced eigen-decomposition for basis reuse—with time complexity $O(m^3 + p m^2)$. Theoretical contributions include a basis optimality verification theorem, local bounds on the objective function, and a sufficient condition guaranteeing successful basis reuse. Experiments demonstrate near cubic-to-quadratic speedups over cold-start reoptimization across instances, substantially reducing computational overhead. The approach is particularly effective in optimization settings requiring repeated solution of similar LP instances.
This study addresses the unclear intrinsic mechanism by which the Local Linear Minimization Oracle (LMO) solves constrained convex optimization without explicit projections. We demonstrate that when the ball radius does not exceed the Polyak radius, a local LMO step is equivalent to an Euclidean projection onto the intersection of a separating hyperplane and the feasible set. Building upon convex optimization theory and Hölder continuous gradient analysis, we construct a unified projection framework indexed by localized halfspace depth. This work provides the first characterization of the implicit projection geometry underlying the local LMO and its theoretical connections to classical methods. Furthermore, we prove that the proposed framework achieves universally optimal convergence rates in the non-accelerated setting, matching the rates of Nesterov’s universal gradient method under specific conditions.
This work addresses the convergence guarantees of stochastic line search optimization for over-parameterized models under interpolation conditions. We establish a necessary and sufficient condition on the search direction—applicable to a broad class of methods—that ensures finite termination and bounded backtracking steps, and rigorously prove linear convergence under the Polyak–Łojasiewicz (PL) assumption. The condition unifies major first-order strategies—including momentum, conjugate gradient, and adaptive preconditioning—providing a verifiable theoretical foundation for their principled integration with stochastic line search. Our analysis fills a critical gap in the convergence theory of stochastic line search methods and significantly extends both the applicability and reliability of efficient first-order optimization in interpolation learning regimes.
Classical iteration complexity analyses for first-order optimization methods rely on global Lipschitz continuity of the gradient, failing to exploit beneficial local smoothness—where the Lipschitz constant varies across regions—and thus incur unnecessary conservatism. Method: We introduce “glocal smoothness,” a novel structural assumption that simultaneously captures both global and local smoothness properties of the objective function—without dependence on algorithmic trajectories—thereby enabling trajectory-agnostic complexity bounds governed solely by intrinsic function constants. Contribution/Results: Under glocal smoothness, we establish improved iteration complexity for gradient descent with backtracking line search—surpassing that of fixed-step accelerated methods. Moreover, we provide a unified, refined convergence analysis for diverse algorithms including Polyak’s step size, adaptive gradient descent (AdGD), coordinate descent, stochastic and deterministic gradient methods, and nonlinear conjugate gradient, yielding significantly tighter complexity bounds across all cases.
This work addresses the challenge of explicitly controlling overfitting during fine-tuning of pretrained Transformers by formulating it as a bilevel optimization regularized framework. It introduces, for the first time, a linear programming–driven local search mechanism that leverages validation gradients and training Hessian information from a warm-up phase to construct a validation-aware descent direction. This enables joint, task-adaptive optimization of both model parameters and regularization hyperparameters without requiring repeated full retraining. Experimental results demonstrate significant reductions in test perplexity on GPT-2 Small and WikiText-2, with particularly pronounced gains in settings prone to overfitting. The approach consistently yields stable improvements across diverse layer configurations and regularization settings.
该研究提出了一种名为CR-P-LPN的方法,通过修复和重用Wolfe子程序的终端围栏来解决线性规划问题,以提高求解效率。
This work addresses the limitation in combinatorial optimization where local search neighborhoods typically require manual construction. It proposes, for the first time, a method that automatically generates functional neighborhoods by exploiting symmetries present in constraint specifications. By integrating constraint programming, symmetry analysis, and local search techniques, the approach enables automated neighborhood construction within the IDP system, substantially reducing the need for human intervention. Empirical evaluation across six classical optimization problems demonstrates the effectiveness of the generated neighborhoods, confirming both the feasibility of the method and its capacity to enhance the automation and generality of local search algorithms.
This work addresses the issue that quadratic penalty relaxations of binary linear programs often yield spurious or infeasible local minima. To overcome this, we propose a class of QUBO relaxation models satisfying specific structural conditions that guarantee all local minima are feasible and strictly binary. Leveraging these conditions, we derive novel differentiable relaxations for classical combinatorial optimization problems—including open-pit mining, the 0–1 knapsack problem, and the traveling salesman problem—and solve them using gradient-based optimizers such as projected gradient descent and Adam. Experimental results demonstrate that the proposed approach reliably converges to valid binary solutions, thereby establishing clear theoretical guarantees and delineating the applicability boundaries of differentiable optimization as a local solver for combinatorial problems.
This work addresses the lack of machine-verifiable formalizations of line search methods in nonlinear optimization, which has hindered algorithmic reliability. Within the Lean 4 theorem prover, it presents the first systematic formalization of several classical line search criteria—including Armijo, Goldstein, Wolfe, and their nonmonotone variants—alongside rigorous definitions of gradient descent, descent directions, and backtracking step-size selection. The study fully verifies the Zoutendijk convergence theorem within this framework, thereby establishing a comprehensive formal foundation for line search theory. This contribution significantly enhances the verifiability and trustworthiness of nonlinear optimization algorithms through mechanized mathematical reasoning.