🤖 AI Summary
This work addresses the challenge of visualizing the implicit periodic phase structures—such as stance and swing phases—embedded in deep reinforcement learning policies for locomotion control. To this end, the authors propose a feature augmentation mechanism that incorporates multi-step temporal information by extending state observations to include the current action, next state, and next action. Combined with a clustering approach that suppresses self-transitions and automatically determines the optimal number of clusters, this method enables the automatic extraction of latent phase structures from policy trajectories. Evaluated on MuJoCo’s Ant-v5, HalfCheetah-v5, and Walker2d-v5 environments, the approach significantly enhances the clarity of identified phases and the regularity of phase transition rules, outperforming existing methods.
📝 Abstract
Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant, and Walker2D. However, visualizing the motion structures internally obtained by a trained policy function implemented as a deep neural network remains challenging. It is known from biomechanics and related fields that locomotion control is realized through the repetition of motion phases such as the stance phase and swing phase. In this study, we propose a framework for uncovering latent motion phase structures from trajectories generated by locomotion control policies through interaction with the environment. The proposed method extends the clustering features from state observations alone to augmented features including actions, next states, and next actions, and introduces a method for determining the number of clusters that suppresses self-transitions. Applying the proposed method to three environments -- Ant-v5, HalfCheetah-v5, and Walker2D-v5 -- we successfully identified phase structures with clearer and more regular transition rules than those obtained by the existing method.