🤖 AI Summary
This study addresses the challenges of representing and performing real-time inference over high-dimensional, heterogeneous action spaces for humanoid robots within discrete Vision-Language-Action (VLA) models. To this end, it proposes Holo-M, the first discrete VLA model tailored for mobile manipulation. Methodologically, action tokens are directly expanded into the language model vocabulary to eliminate knowledge isolation, while a unified factorized action tokenizer is designed to support joint training on multi-source data. During inference, grouped discrete diffusion decoding replaces autoregressive generation, synergizing with specialized tokenizers for end-to-end control, torso, hands, and kinematics to achieve efficient whole-body control. Evaluated on the SIMPLE benchmark, Holo-M achieves state-of-the-art success rates in both generalist and specialist settings, significantly outperforming existing methods.
📝 Abstract
Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.