🤖 AI Summary
This work addresses limitations in existing motion priors for reinforcement learning in humanoid robots, where temporal priors rely on ordered action sequences and pose priors struggle to effectively guide policy-induced pose transitions. The authors propose PFM-HR, a reusable pose flow matching prior that, for the first time, applies flow matching to large-scale unordered pose data to construct a frozen yet transferable prior model. They further introduce the Pose Geometric Score (PGS), which modulates tracking rewards by measuring the consistency between joint changes in policy rollouts and the local geometry of the prior pose manifold, thereby guiding structured exploration. The method significantly improves performance in both single-action and general action-tracking tasks, demonstrating particularly strong gains on highly dynamic motions.
📝 Abstract
Motion priors improve reinforcement learning for physics-based humanoid tracking, but temporal priors require ordered motion clips, while pose priors provide limited guidance for policy-induced pose transitions. We present Pose Flow Matching for Humanoid Robots (PFM-HR), a reusable flow matching prior trained directly on large scale unordered pose data. PFM-HR introduces the Pose Geometry Score (PGS), which quantifies how joint coordinate changes during rollouts align with the local geometry of pose variation captured by the prior. Using PGS to modulate the tracking reward guides policy exploration toward structured pose changes while keeping the prior frozen across tracking tasks. Experiments demonstrate that PFM-HR improves both single motion and general motion tracking, especially for highly dynamic motions.