Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language models typically decouple understanding, dynamics prediction, and action generation into separate architectures. To address this limitation, this work proposes Devol-ONE, a unified autoregressive mixture-of-experts Transformer that integrates multimodal understanding, latent-space dynamics prediction, and action generation within a single framework. The core innovation lies in a layer-wise joint attention mechanism that enables action experts to be continuously guided by the synergistic interplay of semantic reasoning and physical dynamics. Furthermore, the model incorporates a V-JEPA-pretrained dynamics flow to facilitate efficient representation learning. Experimental results demonstrate that the proposed approach achieves superior performance on both the LIBERO benchmark and real-world bimanual robotic manipulation tasks, thereby validating the effectiveness of the unified architecture.
📝 Abstract
Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision language understanding, latent world dynamics prediction, and action generation within a single autoregressive framework. Instead of encoding vision language tokens once and feeding them to the action expert, Devol-ONE runs autoregressive prediction jointly across a vision language stream and a V-JEPA pretrained dynamics stream, attending to the vision language key-value cache at every layer to forecast future latent states under language guidance. The action expert is in turn shaped continuously by semantic reasoning and predicted physical dynamics rather than by a fixed representation computed in advance. Extensive experiments are conducted on LIBERO, LIBERO-PLUS, RoboTwin2.0 along with real-world evaluation on Flexiv single-arm and dual-arm setups. Ablation studies show the effectiveness of dynamic stream prediction and layer-wise unified attention to validate our model architectural coherency.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
World Action Models
latent world dynamics prediction
action generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture of Transformers
Autoregressive framework
Latent world modeling
Vision-Language-Action
V-JEPA
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Hongyi Cai
Hongyi Cai
University of Malaya
Data-centric AIAI for EfficiencyComputer Vision
Y
Yi Herng Ong
Devol Robots, San Francisco, CA, United States
T
Tingshiuan C. Wu
Devol Robots, San Francisco, CA, United States
L
Lim Chiew Hui
Devol Robots, San Francisco, CA, United States
H
Hanxia Li
Devol Robots, San Francisco, CA, United States
K
Kehong Guo
Devol Robots, San Francisco, CA, United States
S
Sze Yuan Cheong
Devol Robots, San Francisco, CA, United States