🤖 AI Summary
This study addresses the absence of a unified benchmark for touch-enabled robotic manipulation by constructing a vision-tactile manipulation benchmark encompassing four capability dimensions with paired simulation-to-real scenarios, systematically evaluating the performance of VLA, WAM, and VTLA policies. Methodologically, this work proposes OpenVTLA, a tactile-augmented framework that integrates optimal representations with ensemble strategies. By leveraging vision-tactile fusion, pretrained model fine-tuning, and cross-domain transfer learning techniques, it thoroughly investigates sim-to-real co-training mechanisms. Ultimately, this project establishes a unified evaluation platform for vision-tactile manipulation, significantly enhancing the reliability and performance of cross-domain policy learning.
📝 Abstract
Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot manipulation across simulation and the real world. To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world. OpenViTac organizes contact-rich manipulation into four tactile-relevant capability dimensions and provides paired simulation-real-world settings for consistent evaluation of VLA, WAM, and VTLA policies. Building upon this benchmark, we investigate how different tactile representations and integration strategies affect the performance of pretrained VLA models. Correspondingly, we introduce OpenVTLA, a tactile augmentation framework that combines the best-performing representation and integration strategy. Furthermore, we leverage the paired benchmark setting to study sim-real co-training and analyze factors affecting cross-domain policy learning. Together, OpenViTac provides a unified platform for evaluating and advancing visuo-tactile robot manipulation.