Score
Designs, builds, or analyzes differentiable continuous relaxations of discrete or structural decision variables by replacing discrete choices with continuous surrogates and constructing differentiable objective and constraint approximations so gradient-based methods (e.g., projected gradient descent) can optimize soft variables. Also develops discretization or projection procedures that restore feasible discrete solutions while preserving constraints and approximation guarantees and accommodating additional requirements such as privacy or query-based validation.
This work addresses the inefficiency of global optimization when standard neural network surrogates are embedded into mixed-integer linear programs (MILPs), a challenge stemming from the lack of control over their structural properties. The authors propose a novel differentiable regularizer that, for the first time, approximates the full gradient of the LP relaxation gap with respect to network parameters, enabling direct optimization of key structural attributes such as big-M constants, the number of unstable neurons, and the LP relaxation gap itself. Built upon ReLU networks and MILP formulations, the method leverages gradients from LP dual variables and requires no custom automatic differentiation. Experiments demonstrate up to four orders of magnitude reduction in MILP solve time on nonconvex benchmark functions and two-stage stochastic programming problems, all while preserving predictive accuracy.
Deep learning struggles to rigorously incorporate hard constraints due to the lack of differentiable, constraint-aware layers. Method: We propose an end-to-end trainable framework that embeds generic convex optimization problems as differentiable layers within neural networks. We establish the first unified differentiability theory for arbitrary differentiable convex optimization, deriving exact gradients via the implicit function theorem and convex analysis, and extend automatic differentiation to support parameterized convex layers. Unlike prior work restricted to quadratic programming, our approach enables rigorous modeling of linear, semidefinite, and general conic constraints. Contribution/Results: Experiments demonstrate substantial improvements in generalization and constraint satisfaction across control, logical reasoning, and physics-guided learning tasks—effectively bridging a critical gap between convex optimization theory and deep learning practice.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
In combinatorial optimization, the empirical risk w.r.t. model parameters is piecewise constant, hindering gradient-based optimization and lacking theoretical generalization guarantees. Method: For contextual stochastic optimization with complex objectives, we propose a perturbation-driven risk smoothing strategy. Our approach integrates statistical learning models with a surrogate combinatorial optimization oracle to construct a context-aware, generalization-controllable decision framework. Contribution/Results: We establish the first unified generalization bound incorporating perturbation bias, statistical error, and optimization error. We introduce the notion of “uniform weak consistency” to characterize the coupled stability between the learning model and the surrogate oracle, proving its universality under mild assumptions. Experiments on stochastic vehicle scheduling demonstrate strong generalization performance. This work provides the first verifiable theoretical generalization framework for contextual stochastic optimization.
Classical $L$-smoothness-based convergence analyses fail for nonsmooth neural networks (e.g., ReLU networks), leading to systematic theoretical misjudgments in nondifferentiable settings. Method: The authors employ nonsmooth optimization theory, Lipschitz analysis, and comparative convergence studies to rigorously characterize the behavior of normalized gradient descent methods (NGDMs). Contributions/Results: They establish that NGDMs exhibit fundamentally distinct convergence dynamics compared to standard gradient descent; reveal that $L_1$ regularization can counterintuitively increase parameter magnitudes—undermining pruning efficacy—in nonsmooth regimes; extend the “Edge of Stability” phenomenon to nonconvex, nonsmooth functions for the first time; and disprove the implicit assumption that algorithms like RMSProp behave identically in differentiable versus nondifferentiable settings. Collectively, these findings challenge the prevailing smoothness-dependent paradigm in deep learning optimization and lay the groundwork for new theoretical frameworks tailored to realistic neural network architectures.
This study addresses challenging 0-1 integer programming problems, particularly large-scale non-convex quadratic knapsack instances. It proposes an autoregressive differentiable optimization framework that leverages a Transformer to sequentially predict variables, thereby strictly maintaining feasibility. By integrating Lagrangian penalties with the Gumbel-softmax strategy, the method efficiently explores the solution space. A core innovation lies in revealing a "quantum tunneling"-like effect, wherein continuous weights traverse potential barriers within the relaxed objective landscape, significantly accelerating global optimization. Empirically, on dense benchmarks comprising tens of thousands of variables, the proposed approach consistently outperforms state-of-the-art open-source solvers in solution quality.
本文提出一种基于张量分解的代理模型方法,通过直接整合可行性信息解决离散黑盒优化问题,有效提高样本效率。
This work addresses stochastic sequential decision-making under hard constraints and combinatorial action spaces, where existing methods struggle to simultaneously ensure scalability and strict feasibility. The authors propose embedding differentiable convex optimization within the policy network: a neural network outputs continuous action targets, which are projected via quadratic programming onto a relaxed feasible set, and dual information is leveraged to map these projections to integer solutions that guarantee constraint satisfaction, enabling end-to-end training. This approach is the first to achieve full coverage of the action space under interactive hard constraints, with provable bounds on integer projection error, thereby combining the expressive power of mixed-integer linear programming (MILP) with the scalability of deep reinforcement learning. Experiments show an average optimality gap below 1% on small instances; on large-scale networks, it outperforms state-of-the-art base-stock policies by up to 9.75% and rolling-horizon stochastic programming by at least 7.7%; in an ASML industrial case study, it reduces costs by up to 3.22%.
This work addresses the challenge in bilevel optimization where non-isolated minima manifolds in the lower-level problem render the upper-level objective nondifferentiable. To overcome this, the authors propose a “select-then-differentiate” framework that, under a local Polyak–Łojasiewicz condition, introduces a unique optimistic selection to ensure hypergradient differentiability and enables explicit hypergradient computation via the pseudoinverse. Theoretically, global uniqueness of the lower-level solution is unnecessary; local smoothness of the upper-level objective is guaranteed as long as the selected solution is nondegenerate on the manifold. The proposed HG-MS method converges to stationary points of the optimistic objective, with complexity governed by the intrinsic dimension of the manifold. Empirically, it significantly outperforms existing approaches on LLM source reweighting tasks, achieving state-of-the-art results on GSM8K and MATH benchmarks and leading performance on MT-Bench.
This work addresses the challenges of credit assignment and high gradient variance in reinforcement learning with hybrid discrete-continuous action spaces, where conventional policy gradient methods suffer significant performance degradation, particularly in high-dimensional continuous action settings. To overcome these limitations, the paper proposes Hybrid Policy Optimization (HPO), which integrates pathwise derivatives and score function gradients to construct an unbiased hybrid gradient estimator within differentiable simulators. Theoretical analysis reveals that the cross-term in the hybrid gradient becomes negligible near discrete optimal responses, justifying an approximately decoupled update strategy that effectively reduces variance. Empirical results demonstrate that HPO substantially outperforms Proximal Policy Optimization (PPO) on inventory control and switched linear quadratic regulator tasks, with performance gains increasing as the dimensionality of the continuous action space grows.