Score
Design and implement policy optimization methods that enforce block-error-rate (BLER) constraints via Lagrangian dual techniques, producing BLER-constrained policies (including Lagrangian-constrained PPO variants) that optimize throughput or other rewards while meeting a specified reliability target. This involves integrating Lagrangian multiplier updates into the training loop to automatically balance objective and constraint satisfaction and avoid manual penalty calibration.
In safe reinforcement learning (SRL), the selection and update of the Lagrange multiplier λ lack theoretical grounding and empirical validation; λ is highly sensitive, and automatic updates often suffer from oscillation, undermining algorithmic stability and performance. To address this, we propose the λ-profile visualization technique to demonstrate that the optimal λ* admits no universal intuitive rule. We further design an adaptive multiplier update framework integrating Lagrangian dual optimization with PID control. Our method is rigorously evaluated across multiple SRL benchmarks. Experiments show that automated λ adaptation not only surpasses performance achieved with a fixed optimal λ but also yields smoother, more stable learning trajectories. While PID control effectively suppresses oscillations, it entails a trade-off between robustness and hyperparameter tuning overhead. The implementation is open-sourced, establishing a new empirical and methodological paradigm for studying constraint optimization stability in SRL.
To address the challenge of joint optimization between constraint embedding and policy learning in constrained reinforcement learning (CRL), this paper establishes a unified equivalence framework bridging CRL and feedback control. Specifically, Lagrange multiplier updates are reformulated as an optimal feedback control problem, and a multiplier-guided policy learning mechanism is introduced to enable end-to-end co-optimization. Theoretically, we show that PID-Lagrangian methods constitute only a special case within this broader framework. Methodologically, we pioneer the integration of model predictive control (MPC) into Lagrangian optimization, proposing Predictive Lagrangian Optimization (PLO)—a novel paradigm for adaptive constraint handling. Evaluated on a multi-task constrained RL benchmark, PLO significantly expands the feasible policy region (+7.2%) while preserving average reward performance, demonstrating its effectiveness, generalizability, and robustness.
Safety-constrained reinforcement learning often suffers from constraint violations and training instability. To address this, we propose Proactive Constraint Policy Optimization (PCPO), which introduces a predictive penalty mechanism that imposes costs before the policy approaches safety boundaries, and a boundary-aware intrinsic reward to guide safe exploration. Theoretically, we establish—for the first time—upper and lower bounds linking the duality gap to policy update performance, thereby guaranteeing convergence and stability. Methodologically, PCPO unifies Lagrangian relaxation, barrier functions, intrinsic rewards, and policy iteration into an end-to-end, safety-driven optimization framework. Experiments demonstrate that, compared to existing post-hoc correction methods, PCPO significantly reduces constraint violation rates while improving robustness and convergence stability across diverse safety-critical benchmarks.
This work addresses training instability in safe reinforcement learning under state-dependent safety constraints, which arises from dual gradient ascent-induced oscillations. To mitigate this issue, the paper proposes the ALaM framework, which introduces, for the first time, a stable training mechanism for state-dependent Lagrange multiplier networks. The approach incorporates an augmented Lagrangian quadratic penalty term to alleviate update delays and enhance local convexity, while employing a supervised regression objective to optimize the multiplier network and improve convergence. Theoretical analysis guarantees convergence of the multipliers and recovery of the optimal policy. Implemented on top of SAC as SAC-ALaM, the method significantly outperforms existing safe RL algorithms across multiple benchmarks, achieving higher safety rates and superior returns while enabling stable training and learning well-calibrated, risk-sensitive multipliers.
To address the low sample efficiency and difficulty in ensuring feasibility during policy training for constrained Markov decision processes (CMDPs) in high-stakes settings, this paper proposes Two-Stage Deep Decision Rules (TS-DDR). TS-DDR is the first method to integrate Lagrangian duality theory into an end-to-end policy training framework, jointly optimizing primal feasibility and dual convergence via deterministic forward solving of the constrained optimization subproblem and closed-form dual gradient backpropagation. By unifying stochastic gradient descent, deterministic optimization solvers, and neural network policy parameterization, it avoids optimality loss induced by convex relaxations. Evaluated on a real-world hydropower scheduling task in Bolivia, TS-DDR achieves significantly higher solution quality while reducing computational time by one to two orders of magnitude compared to state-of-the-art methods.
This study investigates the statistical properties of Lagrange multipliers in constrained maximum likelihood estimation and least squares problems, along with their implications for numerical optimization. Leveraging large-sample theory, it establishes that under correctly specified models, Lagrange multipliers converge in probability to zero as the sample size grows, a result extended to high-dimensional settings such as deep learning. Building on this asymptotic behavior, the work provides the first statistical justification for initializing Lagrange multipliers at zero and integrates this insight into constrained optimization algorithms, including augmented Lagrangian methods and sequential quadratic programming. Numerical experiments demonstrate that this initialization strategy substantially enhances algorithmic stability and convergence efficiency in applications such as constrained regression and dynamic discrete choice models.
This work addresses the challenge of balancing task performance and strict safety constraints in robotic control within safe reinforcement learning. The authors propose PPO-EAL, a novel framework that integrates the exact augmented Lagrangian method into Proximal Policy Optimization (PPO). By combining clipped policy updates with quadratic penalty terms, the approach provides theoretical guarantees for constraint satisfaction, while a momentum-based update rule for Lagrange multipliers enhances dual variable stability. PPO-EAL is the first first-order safe RL method to achieve exact constraint enforcement without requiring excessively large penalty coefficients, offering both convergence guarantees and deployment robustness. Experiments demonstrate superior performance over existing methods across multiple robotic tasks, achieving simultaneous improvements in reward and safety. Notably, the method enables zero-shot sim-to-real transfer for a gear assembly task, significantly boosting success rate and operational robustness.
This work addresses the challenge in safe reinforcement learning where delayed constraint corrections often cause policy oscillations near safety boundaries and prolonged constraint violations. To mitigate these issues, the authors propose Constraint-Sensitive Policy Optimization (CSPO), which incorporates local constraint sensitivity into first-order gradient updates by leveraging the signed shortest distance to guide policy correction. This approach preserves the Karush–Kuhn–Tucker (KKT) solution of the original constrained optimization problem while effectively suppressing oscillations and accelerating safe recovery. Built upon a first-order primal-dual framework, CSPO seamlessly integrates with deep reinforcement learning algorithms. Empirical evaluations on navigation and locomotion tasks demonstrate that CSPO significantly outperforms state-of-the-art methods in terms of safety recovery speed, reward retention, and constraint-satisfaction performance.
This work addresses stochastic sequential decision-making under hard constraints and combinatorial action spaces, where existing methods struggle to simultaneously ensure scalability and strict feasibility. The authors propose embedding differentiable convex optimization within the policy network: a neural network outputs continuous action targets, which are projected via quadratic programming onto a relaxed feasible set, and dual information is leveraged to map these projections to integer solutions that guarantee constraint satisfaction, enabling end-to-end training. This approach is the first to achieve full coverage of the action space under interactive hard constraints, with provable bounds on integer projection error, thereby combining the expressive power of mixed-integer linear programming (MILP) with the scalability of deep reinforcement learning. Experiments show an average optimality gap below 1% on small instances; on large-scale networks, it outperforms state-of-the-art base-stock policies by up to 9.75% and rolling-horizon stochastic programming by at least 7.7%; in an ASML industrial case study, it reduces costs by up to 3.22%.