🤖 AI Summary
This paper addresses the insufficient modeling and enforcement of safety constraints in Safe Reinforcement Learning (SafeRL) and Constrained Markov Decision Processes (CMDPs), particularly in single- and multi-agent settings. To this end, it establishes— for the first time—the unified mathematical framework bridging SafeRL and Safe Multi-Agent RL (SafeMARL). Methodologically, it proposes a systematic approach integrating constrained optimization, policy-gradient-based safety guarantees, safe exploration strategies, multi-agent game-theoretic modeling, and rigorous CMDP-theoretic analysis. The work distills five key open problems in SafeRL/SafeMARL, three of which specifically address novel challenges arising in multi-agent safety-aware cooperation and competition. As a result, it delivers a comprehensive technical guide encompassing formal definitions, theoretical theorems, algorithmic designs, and emerging research directions—constituting the first authoritative reference that simultaneously ensures theoretical rigor and practical applicability for SafeRL and SafeMARL.
📝 Abstract
Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and deployment of agents. This survey provides a mathematically rigorous overview of SafeRL formulations based on Constrained Markov Decision Processes (CMDPs) and extensions to Multi-Agent Safe RL (SafeMARL). We review theoretical foundations of CMDPs, covering definitions, constrained optimization techniques, and fundamental theorems. We then summarize state-of-the-art algorithms in SafeRL for single agents, including policy gradient methods with safety guarantees and safe exploration strategies, as well as recent advances in SafeMARL for cooperative and competitive settings. Additionally, we propose five open research problems to advance the field, with three focusing on SafeMARL. Each problem is described with motivation, key challenges, and related prior work. This survey is intended as a technical guide for researchers interested in SafeRL and SafeMARL, highlighting key concepts, methods, and open future research directions.