Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension
This work addresses the challenge of visualizing the implicit periodic phase structures—such as stance and swing phases—embedded in deep reinforcement learning policies for locomotion control. To this end, the authors propose a feature augmentation mechanism that incorporates multi-step temporal information by extending state observations to include the current action, next state, and next action. Combined with a clustering approach that suppresses self-transitions and automatically determines the optimal number of clusters, this method enables the automatic extraction of latent phase structures from policy trajectories. Evaluated on MuJoCo’s Ant-v5, HalfCheetah-v5, and Walker2d-v5 environments, the approach significantly enhances the clarity of identified phases and the regularity of phase transition rules, outperforming existing methods.