🤖 AI Summary
This study addresses the limitation of vision-only video world models in capturing critical manipulation events such as contact and release, which constrains predictive accuracy. We propose a video world model for dexterous manipulation that integrates tactile signals. Built upon a pretrained video diffusion Transformer with a causal masking mechanism, our approach pioneers the injection of glove-based tactile data into video tokens via zero-initialized residuals, enabling efficient visuo-tactile fusion to enhance future frame prediction. Experimental results demonstrate that the proposed method reduces the hand motion underestimation rate to 9% and decreases perceptual error by 7.4%, significantly outperforming both vision-only approaches and existing visuo-tactile fusion baselines.
📝 Abstract
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future-frame prediction of egocentric manipulation from both observed video and tactile signals from a glove worn on each hand. We condition a pretrained video diffusion transformer on each hand's tactile signal through a zero-initialized residual at the corresponding hand location in the video tokens, while a causal mask prevents predicted frames from accessing future information. Compared with a vision-only model matched in architecture, parameters, and training, DTWM reduces the underestimation of hand motion from 23% to 9%, while reducing the perceptual error in the hand region by 7.4% across three training runs per model. The benefit also increases over the prediction horizon, with the improvement in the later predicted chunks being about 4.1x larger than in the first. DTWM also outperforms other visual-tactile world models under the same setting, and training with touch improves future-frame prediction even when no touch is available at inference. Ablations show that the model benefits from both the magnitude and spatial location of force: replacing the tactile signal with binary contact states, either per hand or per location, increases prediction error. The observed course of the force indicates whether the interaction will persist or change.