🤖 AI Summary
This study addresses the challenge of balancing task performance with safety constraints in safety-critical reinforcement learning. We propose a unified Bellman operator that integrates performance and safety objectives into a joint value function. Convergence is established through two-timescale stochastic approximation and occupation measure-based differential inclusion theory, while neural approximation techniques are incorporated to enable continuous control modeling. Theoretically, this framework ensures that the learned policy maximizes cumulative reward while strictly satisfying trajectory-wide safety constraints. Experimental results demonstrate stable convergence and near-zero safety violations during testing, confirming the algorithm's effectiveness in achieving optimal policy learning under stringent safety requirements.
📝 Abstract
Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.