ฯ„: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision

๐Ÿ“… 2026-07-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of learning effective visuomotor policies under data scarcity while avoiding the high cost of large-scale pretraining. It proposes the ฯ„ framework, which learns action-conditioned, high-dimensional tactile spatiotemporal representations in a latent space through JEPA-inspired future visual supervision, and fuses these with pretrained vision-language features to generate actions. Tactile supervision is introduced only during training, incurring no additional overhead at deployment. Key contributions include the first use of future visual signals to model tactile dynamics, the construction of TacAuraโ€”a multimodal, temporally aligned tactile-vision-language datasetโ€”and a novel tactile-augmented VLA architecture that requires no extra computational cost during inference. Experiments demonstrate that ฯ„ significantly outperforms existing methods across four contact-rich manipulation tasks, exhibiting strong generalization, superior task performance, and robustness.
๐Ÿ“ Abstract
Learning the informative tactile representation while effectively adapting it to pretrained Vision-Language-Action (VLA) models remains challenging at both the data and modeling levels. At the data level, limited task-specific demonstrations constrain representation quality, whereas large-scale pretraining incurs substantial costs. At the modeling level, existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present ฯ„, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation under limited data. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that ฯ„ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness
Problem

Research questions and friction points this paper is trying to address.

tactile representation
Vision-Language-Action models
future visual supervision
contact-rich manipulation
spatiotemporal dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

touch-augmented VLA
future visual supervision
spatiotemporal tactile representation
latent-space learning
TacAura dataset
๐Ÿ”Ž Similar Papers
No similar papers found.