Humanoid World Action Model With Joint State--Action Generation
This study addresses the discrepancy between reference actions and actual execution in humanoid robots by proposing the HWAM model. For the first time, this work introduces post-execution proprioceptive states as an explicit prediction target, jointly generating reference actions and proprioceptive states via a diffusion model. Furthermore, three complementary conditioning pathways are incorporated to connect actions, states, and visual outcomes, enabling multimodal conditional generation for both forward and inverse dynamics. This architecture explicitly models execution discrepancies to optimize action learning. Experimental evaluations conducted on the LimX OLI humanoid robot demonstrate that the proposed method significantly outperforms baseline approaches, achieving a 70.6% success rate in a candy-picking task.