Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of contact-rich interactions in occluded regions for humanoid robots operating under limited visual and proprioceptive feedback. To this end, it proposes integrating whole-body distributed tactile sensing into Vision-Language-Action (VLA) models. Methodologically, predictive objectives are employed to construct multimodal contexts, alongside a predictive supervision mechanism for the tactile pathway that enables latent states to forecast future multimodal representations during action generation, thereby enhancing physical world understanding. Experimental evaluations across five real-world mobile manipulation tasks demonstrate that the proposed approach achieves an average success rate of 75%, outperforming the non-tactile baseline by 43 percentage points. Notably, the method substantially improves control performance in contact-intensive scenarios, such as table sweeping and walking with back-patting interactions.
📝 Abstract
Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions, distributed tactile sensing preserves spatially resolved contact patterns across the robot body. We therefore study how to integrate such whole-body tactile information into vision-language-action (VLA) policies for contact-rich control. Our approach, Uni-VLaT, introduces a tactile pathway whose latent state is trained not only for action generation, but also to predict future tactile, proprioceptive, and visual representations. This predictive objective builds a tactile-anchored multimodal context, encouraging a more structured understanding of the physical world. We evaluate Uni-VLaT on five real-robot tasks covering tactile-triggered locomotion, sustained physical interaction, human-robot contact, and loco-manipulation. Uni-VLaT achieves a 75% average success rate, outperforming a baseline without tactile input by 43 points and a tactile-input baseline without predictive supervision by 7 points. Across two pretrained VLA backbones, our method improves Table Sweeping by 30 points on both backbones and Back-Tap Walking by 85-90 points. Ablations further show that contextualized tactile prediction and absolute future targets are critical to performance. These results indicate that predictive tactile learning provides an effective route for extending pretrained VLA policies to whole-body physical interaction.
Problem

Research questions and friction points this paper is trying to address.

Humanoid Loco-Manipulation
Whole-Body Tactile Sensing
Vision-Language-Action (VLA) Policies
Physical Interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Whole-body tactile sensing
Vision-Language-Action (VLA) policies
Predictive learning
Humanoid loco-manipulation
Multimodal representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zihao Wang
Tsinghua University
S
Shutong Liu
Beihang University
S
Siqi Zheng
Communication University of China
L
Liu Cao
Tsinghua University
R
Ruoqi Chen
Tsinghua University
R
Rundong Liu
Tsinghua University
Yanchao Yang
Yanchao Yang
Assistant Professor, HKU; Stanford University; UCLA
Embodied AIComputer VisionMachine Learning
Mengdi Xu
Mengdi Xu
Stanford University
RoboticsMachine Learning