Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种名为S2C的两阶段强化学习框架,通过分离安全性和竞争性任务的学习来解决机器人在竞争任务中易受攻击的问题。
📝 Abstract
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.
Problem

Research questions and friction points this paper is trying to address.

robot policies
safety
competitive tasks
reinforcement learning
exploitability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety-Filtered Reinforcement Learning
Competitive Task
Adversarial RL
Policy Non-Exploitability