OccluDex: Hierarchical 3D Visuo-Tactile Representation Learning for Egocentric Dexterous Manipulation under Self-Occlusion

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of missing visual evidence caused by manipulator self-occlusion in first-person views, which hinders robust estimation of object geometry and contact states. To this end, we propose a hierarchical 3D visuo-tactile representation learning framework. The method innovatively integrates multi-scale masked autoencoding with cross-modal attention to fuse global geometric structures and local contact information. By pretraining on synchronized human visuo-tactile demonstrations, the perceptual backbone is efficiently transferred to downstream reinforcement learning tasks. Experimental results demonstrate that the proposed framework improves manipulation accuracy on unseen objects by 12.6% in simulation and achieves zero-shot sim-to-real generalization with the Shadow Dexterous Hand in physical experiments.
📝 Abstract
Reliable dexterous manipulation requires continuous estimation of object geometry and hand-object contact throughout interaction. With egocentric sensing, however, the manipulating hand frequently occludes task-relevant object surfaces and contact regions, reducing the visual evidence available for state estimation and thereby making robust closed-loop control and generalization to unseen object geometries particularly challenging. To address this, we present OccluDex, a hierarchical 3D visuo-tactile representation learning framework that integrates global geometric structure with local contact information for robust manipulation under dynamic self-occlusion during hand-object interaction. OccluDex adopts multi-scale masked autoencoding to progressively encode partial 3D geometry and fuses tactile contact tokens with high-level geometric features through cross-modal attention. The encoder is pretrained from synchronized human visuo-tactile demonstrations and transferred as a frozen perceptual backbone for downstream reinforcement learning. We evaluate OccluDex on a faucet rotation task, requiring one full clockwise handle revolution, and a tabletop object reorientation task, requiring a 180-degree tabletop object reorientation without toppling. In simulation experiments, OccluDex demonstrated 12.6% higher accuracy for unseen objects and 8.3% higher accuracy for previously seen objects than the strongest state-of-the-art baseline models. Physical experiments were further performed with a Shadow Hand to demonstrate successful zero-shot sim-to-real generalization on unseen physical objects. This results could enable humanoid egocentric object manipulation for seen and unseen objects even when the manipulating robotic hand occludes vision.
Problem

Research questions and friction points this paper is trying to address.

dexterous manipulation
self-occlusion
egocentric sensing
visuo-tactile representation
state estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

visuo-tactile representation learning
masked autoencoding
cross-modal attention
dexterous manipulation
sim-to-real generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Ziheng Xu
Shanghai Jiao Tong University
Y
Yueyuan Chen
Shanghai Jiao Tong University
X
Xinyuan He
Shanghai Jiao Tong University
G
Guoxing Liu
Shanghai Jiao Tong University
Y
Yuanshuo Tan
Shanghai Jiao Tong University
H
Huiming Pan
Shanghai Jiao Tong University
B
Bin He
Tongji University
S
Shuo Jiang
Tongji University
Peter B. Shull
Peter B. Shull
Professor of Mechanical Engineering, Shanghai Jiao Tong University
Wearable SystemsWearable AIHand Gesture RecognitionHaptic FeedbackIMU