๐ค AI Summary
This work addresses the challenge of learning effective visuomotor policies under data scarcity while avoiding the high cost of large-scale pretraining. It proposes the ฯ framework, which learns action-conditioned, high-dimensional tactile spatiotemporal representations in a latent space through JEPA-inspired future visual supervision, and fuses these with pretrained vision-language features to generate actions. Tactile supervision is introduced only during training, incurring no additional overhead at deployment. Key contributions include the first use of future visual signals to model tactile dynamics, the construction of TacAuraโa multimodal, temporally aligned tactile-vision-language datasetโand a novel tactile-augmented VLA architecture that requires no extra computational cost during inference. Experiments demonstrate that ฯ significantly outperforms existing methods across four contact-rich manipulation tasks, exhibiting strong generalization, superior task performance, and robustness.
๐ Abstract
Learning the informative tactile representation while effectively adapting it to pretrained Vision-Language-Action (VLA) models remains challenging at both the data and modeling levels. At the data level, limited task-specific demonstrations constrain representation quality, whereas large-scale pretraining incurs substantial costs. At the modeling level, existing methods either focus on instantaneous contact states or model temporal interaction dynamics using 6D wrench sequences, leaving high-dimensional tactile signals underexplored. To address these challenges, we present ฯ, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation under limited data. This supervision operates in latent space and is used only during training, adding no deployment overhead. We also introduce TacAura, a dataset of synchronized vision, proprioception, and vision-based tactile signals across four representative contact-rich manipulation tasks. Experiments show that ฯ outperforms existing models and generalizes to unseen objects and scenes, delivering improved manipulation performance and robustness