ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

📅 2026-07-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that contact-rich robotic manipulation relies on tactile cues difficult to capture visually, while real-world tactile data collection is costly and lacks diversity, hindering scalable vision–tactile joint learning. To overcome this, we propose the first world-model-based framework for generating vision–tactile–action trajectories and evaluating policies. By integrating publicly available real tactile data with simulation, we develop an action-conditioned multimodal pretraining model that combines physically grounded tactile modeling with fine-tuning using real-world policies to jointly predict future visual and tactile feedback. Our approach substantially narrows the sim-to-real gap, enables scalable data augmentation and policy validation, and generates physically plausible rollouts in contact-intensive tasks, thereby significantly improving downstream policy performance and enabling accurate evaluation of multimodal outcomes given action sequences.
📝 Abstract
Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust control. However, scaling visuo-tactile robot learning remains difficult because real tactile interaction data are expensive to collect, hardware-dependent, and limited in task and scene diversity. We present ViTacWorld, an action-conditioned visuo-tactile world model for scalable contact-rich robot manipulation. ViTacWorld leverages public real tactile datasets and a constructed simulation environment to scale visuo-tactile-action data, exploiting the fact that tactile signals are directly grounded in physical contact and can exhibit a smaller simulation-to-real gap than purely visual observations. The model is first pretrained with large-scale real and simulated visuo-tactile trajectories, and then finetuned with real-world policy rollouts to better match downstream manipulation behaviors. Given robot actions, ViTacWorld predicts temporally aligned visual observations and tactile feedback, enabling visuo-tactile-action rollout generation. To the best of our knowledge, ViTacWorld is the first framework that uses a world model for robot visuo-tactile-action trajectory generation and policy evaluation. It serves two roles: synthesizing rollouts to improve downstream tactile policies, and evaluating policies by predicting action-conditioned visuo-tactile outcomes under controlled action sequences. Experiments on contact-rich manipulation tasks show that ViTacWorld generates physically meaningful rollouts, improves policy performance through scalable data augmentation, and enables action-conditioned policy evaluation. Project page: https://vitacworld.github.io/
Problem

Research questions and friction points this paper is trying to address.

contact-rich manipulation
tactile sensing
visuo-tactile learning
data scarcity
simulation-to-real gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

visuo-tactile world model
simulation-to-real transfer
contact-rich manipulation
action-conditioned prediction
tactile data augmentation