SpaceDex: Generalizable Dexterous Grasping in Tiered Workspaces

πŸ“… 2026-04-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the challenge of generalizing dexterous grasping under occlusion, narrow gaps, and strong constraints within hierarchical workspaces. The authors propose a hierarchical control framework in which a vision-language model (VLM) is introduced at the high level to perform multi-view spatial reasoning, interpret user intent, and generate target grasp regions. At the low level, an arm–hand feature disentanglement network decouples robotic arm navigation from hand-centric geometric perception for grasp selection, fusing multi-view visual inputs, fingertip tactile feedback, and a small number of corrective demonstrations to enhance robustness. This approach represents the first application of VLMs to multi-view grasp guidance and achieves a 63.0% success rate on over 30 previously unseen objects in real-world experiments, substantially outperforming a desktop baseline (39.0%).

Technology Category

Intelligent Robots: Multimodal Perception & Sensor FusionComputer Vision: Language and VisionMachine Learning: Large Multimodal Models (LMMs)

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
πŸ“ Abstract
Generalizable grasping with high-degree-of-freedom (DoF) dexterous hands remains challenging in tiered workspaces, where occlusion, narrow clearances, and height-dependent constraints are substantially stronger than in open tabletop scenes. Most existing methods are evaluated in relatively unoccluded settings and typically do not explicitly model the distinct control requirements of arm navigation and hand articulation under spatial constraints. We present SpaceDex, a hierarchical framework for dexterous manipulation in constrained 3D environments. At the high level, a Vision-Language Model (VLM) planner parses user intent, reasons about occlusion and height relations across multiple camera views, and generates target bounding boxes for zero-shot segmentation and mask tracking. This stage provides structured spatial guidance for downstream control instead of relying on single-view target selection. At the low level, we introduce an arm-hand Feature Separation Network that decouples global trajectory control for the arm from geometry-aware grasp mode selection for the hand, reducing feature interference between reaching and grasping objectives. The controller further integrates multi-view perception, fingertip tactile sensing, and a small set of recovery demonstrations to improve robustness to partial observability and off-nominal contacts. In 100 real-world trials involving over 30 unseen objects across four categories, SpaceDex achieves a 63.0\% success rate, compared with 39.0\% for a strong tabletop baseline. These results indicate that combining hierarchical spatial planning with arm-hand representation decoupling improves dexterous grasping performance in spatially constrained environments.
Problem

Research questions and friction points this paper is trying to address.

dexterous grasping
tiered workspaces
occlusion
spatial constraints
generalizable manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

dexterous grasping
hierarchical planning
feature separation
multi-view perception
vision-language model
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
W
Wensheng Wang
Sun Yat-sen University, Guangzhou, China.
C
Chuanjun Guo
Stellarobot Company, Shenzhen, China.
W
Wei Wei
Stellarobot Company, Shenzhen, China.
T
Tong Wu
Stellarobot Company, Shenzhen, China.
Ning Tan
Ning Tan
Sun Yat-sen University
RoboticsArtificial Intelligence