OpenViTac: Learning and Benchmarking Visuo-Tactile Policies in a Unified Sim-and-Real Framework

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of a unified benchmark for touch-enabled robotic manipulation by constructing a vision-tactile manipulation benchmark encompassing four capability dimensions with paired simulation-to-real scenarios, systematically evaluating the performance of VLA, WAM, and VTLA policies. Methodologically, this work proposes OpenVTLA, a tactile-augmented framework that integrates optimal representations with ensemble strategies. By leveraging vision-tactile fusion, pretrained model fine-tuning, and cross-domain transfer learning techniques, it thoroughly investigates sim-to-real co-training mechanisms. Ultimately, this project establishes a unified evaluation platform for vision-tactile manipulation, significantly enhancing the reliability and performance of cross-domain policy learning.
📝 Abstract
Tactile feedback provides embodied agents with physical information beyond visual observations, enabling more reliable interaction with the real world. However, despite the rapid progress of vision-tactile-language-action (VTLA) policies, there remains a lack of unified benchmarks for evaluating tactile-enabled robot manipulation across simulation and the real world. To address this gap, we introduce OpenViTac, a visuo-tactile manipulation benchmark for evaluating robot policies across simulation and the real world. OpenViTac organizes contact-rich manipulation into four tactile-relevant capability dimensions and provides paired simulation-real-world settings for consistent evaluation of VLA, WAM, and VTLA policies. Building upon this benchmark, we investigate how different tactile representations and integration strategies affect the performance of pretrained VLA models. Correspondingly, we introduce OpenVTLA, a tactile augmentation framework that combines the best-performing representation and integration strategy. Furthermore, we leverage the paired benchmark setting to study sim-real co-training and analyze factors affecting cross-domain policy learning. Together, OpenViTac provides a unified platform for evaluating and advancing visuo-tactile robot manipulation.
Problem

Research questions and friction points this paper is trying to address.

visuo-tactile manipulation
benchmark
embodied agents
sim-to-real
VTLA policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visuo-Tactile Benchmark
Sim-to-Real Co-training
Tactile Augmentation
VTLA Policies
Tactile Representation
💼 Related Jobs
No related jobs found.
Y
Yifan Wu
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
Q
Qin Li
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI,3Hefei University of Technology
N
Nan Min
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
G
Guojin Zhong
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
H
Haoyu Zhao
4National University of Singapore
Z
Zhiyuan Li
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
H
Houze Xu
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
S
Shengqi Xu
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
X
Xingyao Lin
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
Z
Zijie Diao
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
Zhaoxiang Liu
Zhaoxiang Liu
China Unicom
Computer VisionDeep LearningRoboticsHuman-Computer Interaction
Shiguo Lian
Shiguo Lian
CloudMinds
Shunlin Lu
Shunlin Lu
The Chinese University of Hongkong, Shenzhen
Shihao Zhao
Shihao Zhao
The University of Hong Kong
Generative AIRobust AI
Z
Ziyi Ye
1Fudan University,2Shanghai Key Laboratory of Multimodal Embodied AI
Zuxuan Wu
Zuxuan Wu
Fudan University
Yu-Gang Jiang
Yu-Gang Jiang
Professor, Fudan University. IEEE & IAPR Fellow
Video AnalysisEmbodied AITrustworthy AI