🤖 AI Summary
This study addresses the high inference latency and inadequate responsiveness of vision-language models in dynamic robotic manipulation by proposing a dual-path framework that decouples low-frequency semantic reasoning from high-frequency geometric adaptation. Methodologically, an information interaction module bridges the two pathways to enable online grasp reconstruction and semantic replanning upon failure, while integrating a shape-adaptive network, constraint solving, and real-time RGB-D observations to achieve continuous geometric correspondence updates. Experimental results demonstrate that the proposed approach exhibits superior robustness across six task categories and accelerates geometric adaptation by approximately 46 times compared to conventional verification-based methods.
📝 Abstract
Vision-language models (VLMs) enable open-vocabulary reasoning for robot manipulation, but their high inference latency limits responsiveness in dynamic scenes. Many scene changes, however, alter object geometry without invalidating task intent. We present DualManip, a dual-path framework that decouples infrequent semantic reasoning from responsive geometric adaptation. The semantic path decomposes the task and grounds task-relevant interactions, followed by a constraint-solving module for pose optimization. During execution, the geometric path continuously updates template-to-observation correspondences from live RGB-D observations via a shape-adaptive network. These correspondences transfer task-relevant grasp contacts across observations, enabling online grasp reconstruction under object motion and non-rigid deformation. The Information Interaction Module bridges the two paths by initializing task-relevant grasps from semantic grounding, validating geometric updates, and triggering semantic replanning upon update failures. Real-world evaluation spans six manipulation tasks covering non-rigid deformation, articulated reconfiguration, rigid motion, and high-precision assembly across three settings: static, single-change, and continuous dynamic. DualManip demonstrates superior manipulation robustness, particularly under continuous scene changes, while achieving geometric adaptation approximately 46$\times$ faster than agentic verification and semantic replanning. Our project page: https://lichengxi1.github.io/Dualmanip.