Robust Peak-cost Constrained Reinforcement Learning

πŸ“… 2026-07-16
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of catastrophic risk in safety-critical applications, where traditional constrained Markov decision processes (CMDPs) fail to account for disastrous consequences caused by even a single peak cost violation. We study robust reinforcement learning under peak cost constraints, aiming to maximize expected reward while strictly bounding the maximum cost along any trajectory and accounting for dynamic uncertainty between simulator and real-world environments. We first establish that peak-constrained MDPs may exhibit a non-zero duality gap, thereby challenging the classical duality theory underlying CMDPs. Building on this insight, we propose the first optimization framework that simultaneously ensures robustness and peak cost guarantees, integrating integral probability metric–based robust value estimation with Lagrangian relaxation. Experiments demonstrate that our method enforces constraint violations within a prescribed tolerance Ξ΅ under dynamic perturbations while maintaining strong reward performance.
πŸ“ Abstract
We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most epsilon. Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance.
Problem

Research questions and friction points this paper is trying to address.

robust reinforcement learning
peak-cost constraint
safety-critical
simulator-to-real-world mismatch
constrained MDP
Innovation

Methods, ideas, or system contributions that make the work stand out.

peak-cost constrained RL
robust reinforcement learning
integral probability metrics
duality gap
safety-critical control
πŸ”Ž Similar Papers
No similar papers found.
S
Shilpa Mukhopadhyay
Dept. of Electrical and Computer Engineering, NJIT, Newark, NJ, USA
S
Sourav Ganguly
Dept. of Electrical and Computer Engineering, NJIT, Newark, NJ, USA
S
Santosh Mohan Rajkumar
Department of Mechanical and Aerospace Engineering, The Ohio State University, Columbus, OH, USA
Honghao Wei
Honghao Wei
Assistant Professor of EECS, Washington State University
Reinforcement LearningOptimizationSafe-RL
Debdipta Goswami
Debdipta Goswami
Assistant Professor, The Ohio State University
Non-linear SystemsData-driven EstimationMulti-agent SystemsAutomatic Control
Arnob Ghosh
Arnob Ghosh
Assistant Professor of ECE at New Jersey Institute of Technology
Reinforcement LearningGame thoeryIntelligent Transportation SystemComputer Networks