🤖 AI Summary
This study addresses the lack of interpretability and safety guarantees in reinforcement learning (RL) control for humanoid robots by proposing an RL framework integrated with analytical model-based safety certificates. Leveraging the Angular Momentum Linear Inverted Pendulum (ALIP) model, gait safety certificates are constructed via discrete-time exponential control barrier functions. For the first time, these certificates are simultaneously employed as training signals during policy optimization and as action filters during inference, achieving stable control with minimal intervention. MuJoCo simulations demonstrate that the proposed approach significantly reduces safety violation rates under external disturbances, effectively validating the trade-off between safety assurance and locomotion performance.
📝 Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.