TacDyn-WAM: Learning Implicit Tactile Dynamics in a Heterogeneous Visuo-Tactile World Action Model

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unreliable predictions of tactile world models under deployment drift and the oversensitivity of pixel reconstruction to contact variations by proposing a heterogeneous visuo-tactile world action model. Methodologically, it introduces TacRep, a dynamics-aware tactile representation space, alongside implicit tactile dynamics experts that predict latent dynamics rather than reconstructing observations. Furthermore, a dual-expert architecture integrating masked spatiotemporal prediction, relational structure distillation, and a read-only tactile memory is constructed to enable multi-step prediction and visual interaction within a single forward pass. Experimental results demonstrate that the proposed model achieves an 81.5% success rate on the UniVTAC benchmark and a 71.0% average success rate on real-world robotic tasks, which improves to 85.0% after pretraining, thereby significantly enhancing manipulation robustness.
📝 Abstract
World action models improve robotic manipulation by conditioning actions on predicted futures, yet existing tactile variants largely inherit video-generation pipelines that reconstruct future tactile observations through iterative denoising. Such prediction can become unreliable under deployment drift: small changes in contact position or force may substantially alter tactile pixels even when the underlying contact evolution remains predictable. We introduce TacDyn-WAM, a heterogeneous visuo-tactile world action model that predicts implicit tactile dynamics rather than reconstructing future tactile observations. It learns TacRep, a dynamics-aware tactile target space trained through masked spatio-temporal prediction on tactile clips and regularized by relational structure distillation. A visual expert and an Implicit Tactile Dynamics Expert predict future visual and tactile representations in separate target spaces while interacting through joint attention; the tactile expert predicts future representations and their changes at multiple horizons in a single forward pass, and a read-only tactile memory supplies the current tactile state. On UniVTAC, TacDyn-WAM achieves an average success rate of 81.5% using only the provided demonstrations, reaching state-of-the-art-level performance and remaining competitive with models pretrained on large-scale visuo-tactile trajectories. Ablations confirm the benefits of both tactile pathways and TacRep over pixel-reconstruction and static alternatives. On five real-robot tasks, TacDyn-WAM reaches 71.0% average success, and modest-scale pretraining raises it to 85.0%, further validating our method.
Problem

Research questions and friction points this paper is trying to address.

World Action Model
Tactile Dynamics
Robotic Manipulation
Deployment Drift
Visuo-Tactile Prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Implicit Tactile Dynamics
World Action Model
TacRep
Heterogeneous Visuo-Tactile
Relational Structure Distillation
💼 Related Jobs
No related jobs found.
E
Enyi Wang
Institute for AI Industry Research (AIR), Tsinghua University
M
Mingxin Wang
Institute for AI Industry Research (AIR), Tsinghua University
Q
Quan Shi
Institute for AI Industry Research (AIR), Tsinghua University
H
Hetian Guo
Institute for AI Industry Research (AIR), Tsinghua University
H
Hongyu Wang
Institute for AI Industry Research (AIR), Tsinghua University
X
Xi Wang
Institute for AI Industry Research (AIR), Tsinghua University
Bin Qian
Bin Qian
Post-doctoral researcher at Zhejiang University
internet of thingsedge computingdeep learning
Yupeng Zheng
Yupeng Zheng
Institute of Automation, Chinese Academy of Sciences
Wenxuan Song
Wenxuan Song
The Hong Kong University of Science and Technology (Guangzhou)
Vision-language-action ModelRobotics
H
Houde Liu
Tsinghua University
Yong Xu
Yong Xu
Bio-Computing Research Center, Harbin Institute of Technology, Shenzhen
Image ProcessingPattern RecognitionComputer VisionDeep LearningBiometrics
Cheng Chi
Cheng Chi
Columbia University, Stanford University
robotics
Wenchao Ding
Wenchao Ding
Tenure-track Associate Professor, Fudan University
RoboticsMotion PlanningAutonomous NavigationDecision Making
Y
Yilun Chen
TARS Robotics
Yan Wang
Yan Wang
Tsinghua university; SenseTime
Neural CompressionComputer VisionMachine Learning