🤖 AI Summary
This study addresses the discrepancy between reference actions and actual execution in humanoid robots by proposing the HWAM model. For the first time, this work introduces post-execution proprioceptive states as an explicit prediction target, jointly generating reference actions and proprioceptive states via a diffusion model. Furthermore, three complementary conditioning pathways are incorporated to connect actions, states, and visual outcomes, enabling multimodal conditional generation for both forward and inverse dynamics. This architecture explicitly models execution discrepancies to optimize action learning. Experimental evaluations conducted on the LimX OLI humanoid robot demonstrate that the proposed method significantly outperforms baseline approaches, achieving a 70.6% success rate in a candy-picking task.
📝 Abstract
Humanoid robots are a promising platform for general-purpose manipulation. Recent Vision-Language-Action (VLA) policies learn actions directly from multimodal observations, while World Action Models (WAMs) further incorporate future visual prediction to improve action generation. However, in hierarchical humanoid systems, VLA and WAM policies output reference actions that are subsequently realized through whole-body control, robot dynamics, balance, and contact. This hierarchy creates an action--execution gap: the reference produced by the policy can differ from the motion realized by the robot. Without explicitly modeling the realized body state, future visual prediction must jointly explain scene evolution and discrepancies between reference actions and executed motion, making it difficult to associate an action with its physical outcome. We propose HWAM, a Humanoid World Action Model with joint state--action generation, which makes the robot's post-execution proprioceptive state an explicit prediction target. By jointly generating reference actions and their realized body states, HWAM directly incorporates supervision of executed motion into action learning. HWAM is trained through three complementary conditional paths. The Policy path jointly denoises state--action trajectories conditioned only on current observations, matching deployment conditions. Forward Dynamics Modeling (FDM) predicts future visual observations conditioned on actions and post-execution states, while Inverse Dynamics Modeling (IDM) reconstructs the joint trajectory from visual transitions. Together, these paths connect policy references, realized body motion, and visual outcomes. HWAM achieves the highest success rate among evaluated baselines on three real-robot tasks on the LimX OLI humanoid. On Candy Picking, HWAM achieves a 70.6% success rate, compared with 43.3% for Fast-WAM.