🤖 AI Summary
This study addresses the limitations of existing Vision-Language-Action (VLA) systems that rely on RGB image prediction and re-encoding, which suffer from interface indirection and difficulties in modeling high-dimensional token dynamics. To overcome these bottlenecks, this work proposes constructing an action-conditioned world model directly within the Vision-Language Model (VLM) token space. By compressing visual features into compact state representations, the method enables efficient autoregressive dynamics modeling in a low-dimensional space and directly decodes them into policy actions, thereby eliminating the redundancy of RGB re-encoding. This approach significantly improves open-loop feature fidelity and long-horizon consistency. Furthermore, it achieves a closed-loop simulation performance correlation of 0.794, outperforming the Ctrl-World baseline while maintaining lower inference latency.
📝 Abstract
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World ($r=0.794$ vs.\ $0.583$), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.