A Survey of Safe Reinforcement Learning and Constrained MDPs: A Technical Survey on Single-Agent and Multi-Agent Safety

📅 2025-05-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the insufficient modeling and enforcement of safety constraints in Safe Reinforcement Learning (SafeRL) and Constrained Markov Decision Processes (CMDPs), particularly in single- and multi-agent settings. To this end, it establishes— for the first time—the unified mathematical framework bridging SafeRL and Safe Multi-Agent RL (SafeMARL). Methodologically, it proposes a systematic approach integrating constrained optimization, policy-gradient-based safety guarantees, safe exploration strategies, multi-agent game-theoretic modeling, and rigorous CMDP-theoretic analysis. The work distills five key open problems in SafeRL/SafeMARL, three of which specifically address novel challenges arising in multi-agent safety-aware cooperation and competition. As a result, it delivers a comprehensive technical guide encompassing formal definitions, theoretical theorems, algorithmic designs, and emerging research directions—constituting the first authoritative reference that simultaneously ensures theoretical rigor and practical applicability for SafeRL and SafeMARL.

Technology Category

Multiagent Systems: Multiagent LearningNatural Language Processing: Safety and RobustnessPlanning, Routing, and Scheduling: Planning with Markov Models (MDPs, POMDPs)

Application Category

Security and Privacy: Security and privacy of machine learning and AI applicationsResponsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Safe Reinforcement Learning (SafeRL) is the subfield of reinforcement learning that explicitly deals with safety constraints during the learning and deployment of agents. This survey provides a mathematically rigorous overview of SafeRL formulations based on Constrained Markov Decision Processes (CMDPs) and extensions to Multi-Agent Safe RL (SafeMARL). We review theoretical foundations of CMDPs, covering definitions, constrained optimization techniques, and fundamental theorems. We then summarize state-of-the-art algorithms in SafeRL for single agents, including policy gradient methods with safety guarantees and safe exploration strategies, as well as recent advances in SafeMARL for cooperative and competitive settings. Additionally, we propose five open research problems to advance the field, with three focusing on SafeMARL. Each problem is described with motivation, key challenges, and related prior work. This survey is intended as a technical guide for researchers interested in SafeRL and SafeMARL, highlighting key concepts, methods, and open future research directions.
Problem

Research questions and friction points this paper is trying to address.

Surveying SafeRL and CMDPs for single-agent and multi-agent safety
Reviewing CMDP theory and SafeRL algorithms with safety guarantees
Proposing open research problems in SafeMARL for future advancements
Innovation

Methods, ideas, or system contributions that make the work stand out.

SafeRL based on Constrained MDPs
Policy gradient with safety guarantees
SafeMARL for multi-agent systems
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
International Institute of Information Technology
A
Ankita Kushwaha
International Institute of Information Technology, Hyderabad
P
Pawan Kumar
International Institute of Information Technology, Hyderabad