🤖 AI Summary
Existing vision-language models typically decouple understanding, dynamics prediction, and action generation into separate architectures. To address this limitation, this work proposes Devol-ONE, a unified autoregressive mixture-of-experts Transformer that integrates multimodal understanding, latent-space dynamics prediction, and action generation within a single framework. The core innovation lies in a layer-wise joint attention mechanism that enables action experts to be continuously guided by the synergistic interplay of semantic reasoning and physical dynamics. Furthermore, the model incorporates a V-JEPA-pretrained dynamics flow to facilitate efficient representation learning. Experimental results demonstrate that the proposed approach achieves superior performance on both the LIBERO benchmark and real-world bimanual robotic manipulation tasks, thereby validating the effectiveness of the unified architecture.
📝 Abstract
Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.