Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of interpretability and safety guarantees in reinforcement learning (RL) control for humanoid robots by proposing an RL framework integrated with analytical model-based safety certificates. Leveraging the Angular Momentum Linear Inverted Pendulum (ALIP) model, gait safety certificates are constructed via discrete-time exponential control barrier functions. For the first time, these certificates are simultaneously employed as training signals during policy optimization and as action filters during inference, achieving stable control with minimal intervention. MuJoCo simulations demonstrate that the proposed approach significantly reduces safety violation rates under external disturbances, effectively validating the trade-off between safety assurance and locomotion performance.
📝 Abstract
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.
Problem

Research questions and friction points this paper is trying to address.

Safe Reinforcement Learning
Bipedal Locomotion
Humanoid Robots
Safety Guarantee
Angular Momentum Linear Inverted Pendulum
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safe Reinforcement Learning
Control Barrier Function
Angular Momentum Linear Inverted Pendulum
Bipedal Locomotion
Step-to-Step Prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
V
Victor Paredes
Mechanical and Aerospace Engineering, The Ohio State University, Columbus, OH, USA
Ayonga Hereid
Ayonga Hereid
Assistant Professor, The Ohio State University
RoboticsCyber Physical SystemsOptimal ControlMachine Learning