Score
Design penalty methods: design and implement penalty terms and augmented objective functions that softly enforce constraints by adding quantified penalties to optimization objectives; choose penalty function forms, scale and schedule penalty weights, and shape terms to trade off constraint satisfaction against primary objectives. Analyze how penalty choices affect feasibility, convergence, and solution properties (e.g., collision avoidance or connectivity preservation) and tune them to achieve desired soft-constraint behavior.
Conventional fixed-weight penalty methods in deep learning struggle to simultaneously satisfy constraints and maintain model performance, while suffering from high hyperparameter tuning costs. Method: This paper proposes an end-to-end, constraint-first optimization paradigm. It systematically identifies the fundamental trade-off between constraint strictness and model performance inherent in standard penalty methods and introduces a differentiable augmented Lagrangian framework centered on adaptive Lagrange multipliers—enabling gradient backpropagation and automatic differentiation, and seamlessly integrating with PyTorch and TensorFlow. Contribution/Results: Evaluated across fairness, robustness, and causal constraint tasks, the method achieves 100% constraint satisfaction without compromising classification or regression accuracy, eliminates manual hyperparameter tuning, and improves training efficiency by 3.2×.
This work addresses the challenges of undesired action execution, insufficient policy safety, and poor convergence stability in reinforcement learning. We propose a novel framework integrating structured penalty mechanisms with bidirectional trajectory learning. Methodologically, we introduce the first differentiable structured penalty function coupled with bidirectional (initial- and terminal-state) reinforcement learning, augmented by inverse-dynamics-guided backward sampling and dual-path value function estimation—enabling synergistic forward optimization and backward constraint enforcement in action space. Evaluated on the ManiSkill benchmark, our approach achieves a 92.3% task success rate, outperforming the state-of-the-art by 4 percentage points, accelerating training by 21%, and reducing generalization failure rate by 37%. The framework significantly enhances policy safety, convergence robustness, and sample efficiency.
This paper addresses the $ell_0$-regularized optimization problem under general differentiable loss functions. We propose the first branch-and-bound (B&B) framework applicable to arbitrary losses—not restricted to quadratic—overcoming limitations of prior methods relying on “Big-M” or $ell_2$ relaxations. Our key methodological contribution is a unified, flexible relaxation theory encompassing multiple relaxation strategies, enabling closed-form expressions for all critical B&B quantities: dual bounds, branching rules, and node pruning conditions. Based on this theory, we develop El0ps, an open-source solver supporting user-defined losses and regularizers for plug-and-play $ell_0$ modeling. Experiments demonstrate state-of-the-art performance on classical benchmarks; notably, El0ps is the first to provably solve large-scale and non-quadratic $ell_0$ problems, substantially expanding the computational tractability frontier of $ell_0$ optimization.
In constrained reinforcement learning for continuous control, existing methods struggle with the reward-safety trade-off and suffer from training instability near constraint boundaries. To address these challenges, this paper proposes IP3O—a novel constrained policy optimization algorithm that integrates an adaptive incentive mechanism with an incremental penalty strategy to actively guide safe actions within the constraint critical region, while incorporating a theoretical error bound analysis to ensure robustness. IP3O is the first method to unify dynamic incentive shaping, progressive constraint penalization, and a worst-case optimality error upper bound of $O(sqrt{T})$ within the proximal policy optimization (PPO) framework. Empirical evaluation on benchmark environments including Safety Gym demonstrates that IP3O significantly outperforms state-of-the-art safe RL algorithms: it achieves superior policy performance while strictly satisfying safety constraints and markedly improving training stability.
Existing penalty-based methods for bilevel optimization (BLO) with coupling constraints suffer from high computational overhead due to inner-loop iterations and necessitate small outer-loop step sizes owing to stringent smoothness requirements, leading to poor convergence complexity. Method: This paper proposes a novel penalty reformulation that decouples upper- and lower-level variables, substantially reducing reliance on objective function smoothness. Building upon this, we design PBGD-Free—a single-loop algorithm that eliminates inner-loop optimization and replaces the standard Lipschitz gradient assumption with a curvature condition, enabling smaller penalty coefficients and tighter gradient control. Contribution/Results: We establish theoretical convergence guarantees under mild assumptions. Empirical evaluation on SVM hyperparameter tuning and large-model fine-tuning demonstrates significant reductions in iteration complexity and substantial improvements in training efficiency.
This work addresses the limitations of traditional numerical methods—which rely heavily on gradients and initial guesses—and the slow convergence of evolutionary algorithms in high-dimensional constrained optimization. To overcome these challenges, the paper proposes embedding a population-based stochastic optimizer, such as CMA-ES, into an augmented Lagrangian (AL) framework, replacing local solvers in AL subproblems with gradient-free global search. This approach represents the first systematic integration of the AL method’s robust constraint-handling capabilities with the strong exploratory power of evolutionary algorithms, effectively balancing feasibility enforcement and global exploration. Experimental results demonstrate that the proposed method significantly outperforms both pure evolutionary algorithms and state-of-the-art solvers like IPOPT on standard benchmark problems, particularly excelling in high-dimensional nonconvex landscapes riddled with numerous local minima and saddle points.
This study investigates the statistical properties of Lagrange multipliers in constrained maximum likelihood estimation and least squares problems, along with their implications for numerical optimization. Leveraging large-sample theory, it establishes that under correctly specified models, Lagrange multipliers converge in probability to zero as the sample size grows, a result extended to high-dimensional settings such as deep learning. Building on this asymptotic behavior, the work provides the first statistical justification for initializing Lagrange multipliers at zero and integrates this insight into constrained optimization algorithms, including augmented Lagrangian methods and sequential quadratic programming. Numerical experiments demonstrate that this initialization strategy substantially enhances algorithmic stability and convergence efficiency in applications such as constrained regression and dynamic discrete choice models.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This study addresses the challenge of constructing D-optimal experimental designs for generalized linear models (GLMs) with mixed discrete-continuous factors, where the Fisher information matrix depends on unknown parameters and lacks closed-form solutions, rendering efficient computation difficult. To overcome this, the authors propose a penalty-based particle swarm optimization method (p-PSO) that reformulates the constrained optimization problem into an unconstrained one, enabling direct application of off-the-shelf PSO algorithms. The proposed penalty scheme is algorithm-agnostic and readily extensible to various black-box optimizers. Empirical results demonstrate that the method achieves superior computational efficiency and robustness, successfully generating D-optimal designs for GLMs with mixed factors and confirming the broad applicability of the penalty mechanism across diverse constrained optimization tasks.
This work investigates the problem of simultaneously minimizing static regret and cumulative constraint violation (CCV) in constrained online convex optimization. Focusing on the classical Online Gradient Descent with Projection (OGD+Projection) algorithm, the study establishes—for the first time—a theoretical lower bound on CCV of $\Omega(T^{(d-1)/(2d)})$ for any dimension $d$. This result closes a significant gap in the existing theoretical analysis and reveals fundamental performance limitations of the algorithm in high-dimensional settings. By precisely characterizing this trade-off, the paper provides crucial theoretical insights into the inherent tension between regret minimization and constraint satisfaction in constrained online learning.