🤖 AI Summary
This study addresses the challenge of generalizing world models across diverse robotic embodiments caused by morphological heterogeneity. To this end, it proposes the LAC-WM framework, which pioneers a unified latent action space to overcome representational fragmentation induced by explicit action labels. By employing a latent-action-conditioned architecture that integrates deep reinforcement learning with multimodal data pretraining, the method achieves positive scaling as the number of pretrained embodiments increases. Experimental evaluations demonstrate that LAC-WM improves performance by 46.7% on dexterous manipulation tasks and 11.7% on the LIBERO benchmark. These results indicate that the proposed approach significantly enhances adaptability to unseen embodiments and cross-embodiment learning efficiency.
📝 Abstract
The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This unified action space improves the world model's performance when adapted to previously unseen robot embodiments. We compare LAC-WM with an Explicit Action-Conditioned World Model (EAC-WM), which conditions on explicit motion labels. Our results show that explicit action conditioning leads to disjoint action representations across embodiments, limiting downstream performance when adapting to new robots. We evaluate both models on dexterous manipulation tasks and a modified LIBERO benchmark. LAC-WM improves downstream performance over EAC-WM by up to 46.7% on dexterous manipulation and 11.7% on LIBERO. Crucially, the unified latent action space allows LAC-WM's downstream performance to scale positively with the number of embodiments used during pretraining. In contrast, the disjoint action space in EAC-WM leads to decreased performance as the number of pretraining embodiments increases. These results highlight the importance of a unified action space for efficient cross-embodiment learning, addressing a key challenge in robotics. Project website: https://lacwm.github.io/