🤖 AI Summary
This study addresses the unreliable predictions of tactile world models under deployment drift and the oversensitivity of pixel reconstruction to contact variations by proposing a heterogeneous visuo-tactile world action model. Methodologically, it introduces TacRep, a dynamics-aware tactile representation space, alongside implicit tactile dynamics experts that predict latent dynamics rather than reconstructing observations. Furthermore, a dual-expert architecture integrating masked spatiotemporal prediction, relational structure distillation, and a read-only tactile memory is constructed to enable multi-step prediction and visual interaction within a single forward pass. Experimental results demonstrate that the proposed model achieves an 81.5% success rate on the UniVTAC benchmark and a 71.0% average success rate on real-world robotic tasks, which improves to 85.0% after pretraining, thereby significantly enhancing manipulation robustness.
📝 Abstract
World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.