🤖 AI Summary
This study addresses the challenge of missing visual evidence caused by manipulator self-occlusion in first-person views, which hinders robust estimation of object geometry and contact states. To this end, we propose a hierarchical 3D visuo-tactile representation learning framework. The method innovatively integrates multi-scale masked autoencoding with cross-modal attention to fuse global geometric structures and local contact information. By pretraining on synchronized human visuo-tactile demonstrations, the perceptual backbone is efficiently transferred to downstream reinforcement learning tasks. Experimental results demonstrate that the proposed framework improves manipulation accuracy on unseen objects by 12.6% in simulation and achieves zero-shot sim-to-real generalization with the Shadow Dexterous Hand in physical experiments.
📝 Abstract
Reliable dexterous manipulation requires continuous estimation of object geometry and hand-object contact throughout interaction. With egocentric sensing, however, the manipulating hand frequently occludes task-relevant object surfaces and contact regions, reducing the visual evidence available for state estimation and thereby making robust closed-loop control and generalization to unseen object geometries particularly challenging. To address this, we present OccluDex, a hierarchical 3D visuo-tactile representation learning framework that integrates global geometric structure with local contact information for robust manipulation under dynamic self-occlusion during hand-object interaction. OccluDex adopts multi-scale masked autoencoding to progressively encode partial 3D geometry and fuses tactile contact tokens with high-level geometric features through cross-modal attention. The encoder is pretrained from synchronized human visuo-tactile demonstrations and transferred as a frozen perceptual backbone for downstream reinforcement learning. We evaluate OccluDex on a faucet rotation task, requiring one full clockwise handle revolution, and a tabletop object reorientation task, requiring a 180-degree tabletop object reorientation without toppling. In simulation experiments, OccluDex demonstrated 12.6% higher accuracy for unseen objects and 8.3% higher accuracy for previously seen objects than the strongest state-of-the-art baseline models. Physical experiments were further performed with a Shadow Hand to demonstrate successful zero-shot sim-to-real generalization on unseen physical objects. This results could enable humanoid egocentric object manipulation for seen and unseen objects even when the manipulating robotic hand occludes vision.