Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing Vision-Language-Action (VLA) systems that rely on RGB image prediction and re-encoding, which suffer from interface indirection and difficulties in modeling high-dimensional token dynamics. To overcome these bottlenecks, this work proposes constructing an action-conditioned world model directly within the Vision-Language Model (VLM) token space. By compressing visual features into compact state representations, the method enables efficient autoregressive dynamics modeling in a low-dimensional space and directly decodes them into policy actions, thereby eliminating the redundancy of RGB re-encoding. This approach significantly improves open-loop feature fidelity and long-horizon consistency. Furthermore, it achieves a closed-loop simulation performance correlation of 0.794, outperforming the Ctrl-World baseline while maintaining lower inference latency.
📝 Abstract
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World ($r=0.794$ vs.\ $0.583$), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.
Problem

Research questions and friction points this paper is trying to address.

World Modeling
Vision-Language-Action
Robot Manipulation
Token Space
Autoregressive Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model
Vision-Language Model
Token Space
Robot Manipulation
Compact Representation
💼 Related Jobs
No related jobs found.
C
Chuyao Fu
Southern University of Science and Technology
Xiaowei Chi
Xiaowei Chi
The Hong Kong University of Science and Technology
Multimodal GenerationRoboticsComputer Vision
Y
Yuhan Rui
Southern University of Science and Technology
Y
Yu-kai Wang
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Zezhong Qian
Zezhong Qian
XianJiaotongUniversity
World ModelAutonomous DrivingVideo GenerationRobot Manipulation
Xiaojie Zhang
Xiaojie Zhang
City University of New York, The Graduate Center
Edge computing
Y
Yunfan Lou
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Kevin Zhang
Kevin Zhang
Peking University
ML
Kuangzhi Ge
Kuangzhi Ge
Peking University
Multimodal LearningEmbodied AI
C
Chak Wing Mak
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
Z
Zhiyang Chen
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
A
Athena Zhuoming Zhong
University of Pennsylvania
Hongyang Chen
Hongyang Chen
SUN YAT-SEN UNIVERSITY
SDNCloud ComputingMicroserviceAIOps
Haoran Li
Haoran Li
Institute of Automation,Chinese Academy of Sciences
Artificial IntelligenceRoboticsReinforcement LearningEmbodied Intelligence
Y
Yike Guo
Hong Kong University of Science and Technology
Sirui Han
Sirui Han
The Hong Kong University of Science and Technology
Large Language ModelInterdisciplinary Artificial Intelligence
Shanghang Zhang
Shanghang Zhang
Peking University
Embodied AIFoundation Models