🤖 AI Summary
This study addresses the scarcity of training data and the limited precision of generative models for dexterous grasping in cluttered scenes by constructing the first million-scale 3D scene benchmark and proposing the OmniDex model. Methodologically, high-quality grasp candidates are efficiently selected through 3D object curation and seed filtering strategies. Furthermore, Soft Winner-Takes-All (SWTA) learning is integrated with physics-constrained training to eliminate post-optimization latency, thereby enabling efficient end-to-end inference. Experimental results demonstrate that the proposed approach achieves state-of-the-art performance across diverse scenes, multiple viewpoints, and unseen objects, exhibiting exceptional generalization capability and robustness.
📝 Abstract
Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.