🤖 AI Summary
This study addresses the limitation that safety filtering imposes on policy exploration in reinforcement learning within dynamic obstacle environments by proposing the DODGER framework. Departing from conventional approaches that restrict execution to safe actions, this method eliminates runtime safety filtering and directly executes policy actions during training. By integrating Control Barrier Function (CBF) reference signals with violation penalties, it guides the agent to autonomously acquire obstacle-avoidance behaviors. Coupled with LiDAR perception, the framework achieves safe goal-directed navigation across multiple dynamic obstacles in both simulation and real-world humanoid robot platforms. Consequently, this work provides an effective paradigm for filter-free safe reinforcement learning in dynamic environments.
📝 Abstract
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.