🤖 AI Summary
This study addresses the challenge of diverse encoding architectures and the absence of unified evaluation standards for vision-based tactile sensors. We systematically investigate the performance of tactile encoders and multimodal fusion strategies in contact-rich manipulation tasks. Through over 2,000 controlled comparative experiments conducted in real-world settings—the first of this scale—we comprehensively evaluate computer vision-based tactile encoding alongside end-to-end manipulation policies. Our results demonstrate that no universally optimal solution exists; rather, the most effective encoder and fusion design choices are highly task-dependent. By providing extensive empirical evidence from rigorous physical experimentation, this work offers critical insights to guide the principled design of tactile perception systems for robotic manipulation.
📝 Abstract
Tactile information is essential for contact-rich manipulation tasks in robotics. Vision-based tactile sensors make it particularly easy to design end-to-end manipulation policies with tactile sensing, as they enable the use of existing encoders from computer vision. However, this has led to a huge variety of architectures, training datasets, and evaluation protocols, making it difficult to determine which design choices best encode touch. In this work, we address this gap and present a comprehensive study of tactile encoders and fusion strategies across various contact-rich manipulation tasks in real-world experiments. To enable a controlled comparison, we train and evaluate all models under the same pipeline and experimental setup, comprising more than 2000 real-world rollouts. Our results go beyond other studies that only compare simulation performance, which does not necessarily translate to real-world settings, where large-scale evaluations are needed to obtain reliable statistics. Our key finding is that there is no universally optimal representation or fusion strategy for encoding visual-tactile. Instead, the best encoder backbone and fusion scheme depend strongly on the task.