🤖 AI Summary
This study addresses the strong coupling between obstacle avoidance and balance control in underactuated bipedal robots by proposing a hierarchical reinforcement learning framework. The high-level policy generates velocity commands while the low-level policy executes joint control, incorporating task-agnostic convergent gaits to seamlessly integrate classical motion planners. Policies are jointly trained using the Soft Actor-Critic (SAC) algorithm and validated in PyBullet simulations against A* and RRT* baselines. Results demonstrate that the proposed method achieves success rates of 98% and 88% in static and dynamic environments, respectively. It significantly outperforms conventional hybrid approaches while maintaining near-optimal path efficiency, effectively overcoming existing performance bottlenecks in bipedal locomotion planning and control.
📝 Abstract
A bipedal robot cannot deviate from its path to avoid an obstacle without disturbing its balance, and this coupling is most severe on underactuated platforms such as the biped considered here, which has four actuated joints per leg and no hip or ankle roll. This paper presents a Hierarchical Reinforcement Learning (HRL) framework in which a High-Level (HL) policy observes the robot pose, 36 raycast proximity measurements, moving-obstacle states, and a receding-horizon local goal, and outputs a body-velocity command $(v_x, v_y, \omega_{yaw})$ every ten control steps, while a velocity-conditioned Low-Level (LL) policy tracks each command through PD-controlled joint targets. Both policies are trained jointly with Soft Actor-Critic (SAC). Because the converged gait is task-agnostic, it is frozen and driven by classical planners over the same command interface, yielding three controlled baselines: SAC+A*, SAC+RRT*, and SAC+APF. Across 100 evaluation trials per method in randomized PyBullet environments, the proposed method reaches the goal in 98.0% of static and 88.0% of dynamic trials, against at most 78.0% and 68.0% for the planner hybrids, with path lengths within 4% of the A* reference, and ablations confirm that each observation channel and reward term contributes materially to this performance.