π€ AI Summary
This work addresses two key limitations of existing world action models: the absence of cross-task skill reuse mechanisms and background redundancy that interferes with critical visual information extraction. To this end, we propose a skill reuse framework based on an action experience dictionary and shared action embeddings. Specifically, a pretrained action tokenizer encodes historical trajectories, while a cross-attention mechanism integrates visual conditions with action intent to enable precise prediction. Furthermore, a motion-aware transition loss is introduced to effectively suppress background noise. Extensive evaluations demonstrate the superiority and generalization capability of the proposed method across both simulation benchmarks and real-world cross-embodiment scenarios.
π Abstract
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.