🤖 AI Summary
This study addresses the distribution mismatch and safety hazards arising from the independent training of tracking policies and safety filters for humanoid robots. To overcome these limitations, we propose CoFiT, a method that optimizes pretrained trackers through perception-aware filter fine-tuning. By elucidating the interaction mechanisms between policies and filters, this work establishes design principles for integrating learned trackers with runtime safety filters to achieve safe whole-body motion control. Experimental evaluations demonstrate that the proposed approach reduces simulation violation time by 91%. Furthermore, real-world deployment operates without human intervention while achieving an 83% reduction in safety violations, confirming its effectiveness in bridging the sim-to-real gap for robust humanoid control.
📝 Abstract
Safe whole-body motion is essential for deploying humanoid robots in unstructured environments. Modern humanoid control commonly separates reference specification from execution, with a planner, teleoperator, or motion generator providing a reference that a reinforcement-learning policy tracks through dynamically feasible whole-body control. Runtime safety filters, such as control barrier functions (CBFs), offer a promising approach for enforcing newly introduced constraints via interventions on the tracker's outputs. We show, however, that treating the tracking policy and safety filter independently induces fundamental mismatches, as filtering alters both the executed actions and the induced state distribution. We study this policy-filter interface through case studies that isolate dynamics, objective, and information mismatches, highlight their root causes, and use these insights to develop CoFiT (Constrained Filter-aware Tuning), a filter-aware fine-tuning method for pretrained trackers. Across diverse constraint scenes, CoFiT reduces violation time relative to filter-only training by 91% on TWIST2 and 21% on SONIC, while requiring smaller safety filter corrections. On Unitree G1 hardware, CoFiT reduces violation time by 83% for TWIST2 and completes every trial without operator intervention, whereas 50% of baseline trials require an operator stop. Together, these results provide actionable insights into policy-filter interactions and establish design principles for integrating learned trackers with runtime safety filters.