🤖 AI Summary
This study addresses the data and computational bottlenecks arising from integrating tactile sensing into large-scale pretraining for Vision-Language-Action (VLA) models by proposing SimpleTouch. Built upon the π0.5 architecture, this method eliminates complex multi-stage alignment pipelines by freezing the tactile encoder and incorporating multi-horizon latent action prediction, enabling contact-rich manipulation through single-stage action-supervised fine-tuning. Experimental results demonstrate that SimpleTouch achieves an average success rate of 77.5% on the UniVTAC benchmark using only 50 demonstrations—outperforming FTP-π0.5 by 32.3%—and attains a 71.3% success rate in real-world tasks. These findings significantly surpass existing baselines, validating the efficiency of directly fine-tuning pretrained representations for tactile-dependent robotic manipulation without extensive retraining.
📝 Abstract
Tactile sensing provides essential contact information for robotic manipulation, yet incorporating it into pretrained vision-language-action (VLA) models remains challenging. A common concern is that simply introducing touch during task-specific fine-tuning may fail to bridge the cross-modal gap, yielding limited gains or even reduced success. Consequently, existing methods often rely on large-scale tactile policy pretraining or separate visuotactile alignment, adding data requirements and training stages. We introduce SimpleTouch, a simple VLA extension that augments $π_{0.5}$ with a tactile expert, to test whether these additional stages are necessary. Leveraging all tokens from a frozen pretrained tactile encoder, the expert learns from action supervision and multi-horizon prediction of future tactile latents. This single-stage training uses only task demonstrations, without additional tactile policy pretraining or separate alignment. With 50 demonstrations per task, SimpleTouch achieves the highest success rate among evaluated methods on all six UniVTAC tasks. Its average success rate reaches 77.5%, compared with 45.2% for FTP-$π_{0.5}$ and 66.7% for FTP-1, corresponding to gains of 32.3 and 10.8 percentage points, respectively. Across four real-world tasks, it averages 71.3%, exceeding FTP-1 by 8.8 percentage points. These results demonstrate that, given pretrained VLA and tactile representations, additional tactile policy pretraining is not a prerequisite for strong performance on these tasks, offering a simpler route to contact-rich manipulation. Project page: https://simpletouch-robot.github.io/