Score
Design and implement layers that embed parametrized convex optimization problems as differentiable components in computational graphs, producing primal/dual solutions and gradients with respect to parameters. Ensure the layer adheres to disciplined parametrized programming (DPP) conventions so convexity is preserved, solvers are numerically stable, and differentiation through re-optimization is correct and efficient.
Deep learning struggles to rigorously incorporate hard constraints due to the lack of differentiable, constraint-aware layers. Method: We propose an end-to-end trainable framework that embeds generic convex optimization problems as differentiable layers within neural networks. We establish the first unified differentiability theory for arbitrary differentiable convex optimization, deriving exact gradients via the implicit function theorem and convex analysis, and extend automatic differentiation to support parameterized convex layers. Unlike prior work restricted to quadratic programming, our approach enables rigorous modeling of linear, semidefinite, and general conic constraints. Contribution/Results: Experiments demonstrate substantial improvements in generalization and constraint satisfaction across control, logical reasoning, and physics-guided learning tasks—effectively bridging a critical gap between convex optimization theory and deep learning practice.
Differentiable optimization in robot visual localization often suffers from local minima and gradient distortion—especially in low-light keypoint detection—compromising robustness and accuracy. Method: This paper proposes a certifiably differentiable framework based on polynomial optimization (POP), which reformulates POP problems into semidefinite programming (SDP) relaxations with certified backward propagation. It integrates implicit differentiation with PyTorch-based end-to-end training, ensuring global optimality guarantees while maintaining computational efficiency. Contribution/Results: To the best of our knowledge, this is the first work to enable certified backpropagation through SDP relaxations of POP problems, theoretically guaranteeing gradient correctness. Experiments demonstrate that the method substantially mitigates failure modes of mainstream differentiable optimizers, significantly improving both keypoint detection robustness and localization accuracy under low-light conditions in robotic visual localization tasks.
Existing differentiable quadratic programming (QP) methods rely on solver-specific implementations, hindering seamless integration into neural networks or bilevel optimization pipelines and restricting solver choice. This paper introduces dQP—the first explicit differentiation framework based on the active set, enabling end-to-end differentiability for arbitrary black-box QP solvers (compatible with 15+ mainstream solvers) without modifying solver source code; only the optimal solution and active constraint set are required for backward propagation. dQP unifies convex optimization theory, implicit function differentiation, and automatic differentiation, supporting both dense small-scale and sparse large-scale QP problems. Experiments demonstrate that dQP consistently outperforms prior differentiable QP methods across diverse benchmarks. Moreover, it enables a novel bilevel geometric optimization task, showcasing broad applicability beyond standard QP settings.
Verifying geodesic convexity on Hadamard manifolds—such as the manifold of symmetric positive-definite matrices—is inherently challenging and typically requires case-specific analysis, hindering reliable modeling and optimization in non-Euclidean settings. Method: This paper introduces the *Normalized Geodesic Convex Programming* (DGCP) framework—the first systematic framework for geodesic convexity verification and optimization on Cartan–Hadamard manifolds. It defines geodesic convex atomic functions and composition rules preserving geodesic convexity, grounded in differential geometry and convex analysis, enabling symbolic-level automatic certification. Implemented in Julia as the SymbolicAnalysis.jl library, DGCP seamlessly integrates with manifold optimization solvers for end-to-end modeling and solving. Contribution/Results: DGCP significantly improves modeling reliability and computational efficiency for statistical estimation, matrix learning, and other non-Euclidean optimization tasks, establishing a verifiable convexity foundation for Riemannian optimization.
This work addresses discrete-time nonlinear optimal control problems by unifying classical algorithms—including gradient descent, Gauss–Newton, Newton’s method, and differential dynamic programming (DDP)—within a differentiable programming framework. Methodologically, it introduces the first modular, end-to-end differentiable algorithm template library built upon linear/quadratic approximations (e.g., LQR), enabled by automatic differentiation. Theoretically, it provides a unified derivation of computational complexity and sufficient optimality conditions across all methods. Practically, it incorporates adaptive line search and regularization strategies, and validates efficacy on benchmark tasks such as autonomous racing with a bicycle model. All implementations are open-sourced, demonstrating both efficient gradient propagation and strong generalization across diverse control problems.
This study addresses the challenge of efficiently differentiating through general conic programming embedded learning systems by proposing a solver-agnostic differentiable framework. The method geometrically reduces conic problems to quadratic programs for gradient computation while preserving reference solutions and first-order sensitivity information, ensuring well-definedness at singular points with only a single linear solve required. Explicit gradient formulas are derived for convex nonlinear programs (NLPs), quadratic programs (QPs), second-order cone programs (SOCPs), and semidefinite programs (SDPs), leveraging symmetric linear solvers to enhance computational efficiency. Experimental results validate the correctness of the computed gradients and demonstrate the scalability of backpropagation, achieving significant acceleration on large-scale problems.
This work addresses the critical challenge that modern GPU-accelerated linear programming solvers—such as cuPDLP, which is based on the primal-dual hybrid gradient (PDHG) algorithm—exhibit performance highly sensitive to hyperparameters, yet lack tuning methods with provable generalization guarantees. For the first time, this study establishes structural relationships between hyperparameters and solution trajectories for multiple adaptive techniques in complex first-order LP solvers, including preconditioning, restart strategies, and smoothed weight updates. By integrating convergence analysis of PDHG with a model of structural sensitivity, the authors propose a data-driven hyperparameter learning framework that offers theoretical generalization guarantees under polynomial sample complexity. Experimental results demonstrate that the framework significantly enhances solver efficiency across diverse problem instances.
This work addresses the challenge of enforcing strict output fairness constraints in deep learning models under streaming prediction and small-batch settings, where conventional batch-based fairness methods fall short. To this end, the authors propose a differentiable “fairness layer” integrated at the network output to hard-enforce prescribed fairness criteria. They further introduce the first online primal-dual inference algorithm capable of operating on arbitrarily small batches, thereby overcoming the limitations of traditional batch-constrained approaches. Notably, this is the first method to employ a differentiable optimization layer for enforcing aggregate fairness, complemented by stability analysis within backpropagation to ensure differentiability and convergence during training. Experiments demonstrate that the proposed framework rigorously satisfies fairness requirements without compromising model performance, with both theoretical analysis and empirical results confirming its efficacy.
This work addresses the computational challenges of embedding neural networks into mathematical optimization, where conventional feedforward neural networks (FNNs) yield mixed-integer programming (MIP) reformulations that are computationally expensive and suffer from loose relaxations. To overcome these limitations, the paper proposes using input convex neural networks (ICNNs) as surrogate models, leveraging their inherent convexity to construct tight linear programming (LP) relaxations. The authors establish, for the first time, an exact convex hull-based continuous relaxation of ICNNs over box domains, yielding an LP representation free of integrality gaps. Furthermore, they introduce a novel branch-and-bound algorithm that branches directly on input variables. Demonstrated across applications in humanitarian food aid allocation, oil well trajectory planning, and wine blending, the approach achieves approximation accuracy comparable to FNNs while significantly improving solution speed and scalability.
本文提出了一种符号框架DBLP,用于以高阶可读方式指定和解决双层优化问题,并通过自动正则化和连续逼近法求解。