🤖 AI Summary
This study addresses the disconnection among visual grounding, path planning, and execution decision-making in zero-shot robotic manipulation by proposing a keypoint-enhanced hierarchical architecture. Utilizing semantic 3D representations as a bridge, this method unifies visual grounding and language-conditioned planning through point-level interfaces. Specifically, it employs PointVLM for visual localization and integrates depth information to construct 3D representations, leverages 3DLLM for waypoint planning, and introduces a hybrid grasping module to achieve end-to-end control, all without requiring task-specific demonstration training. Experimental evaluations across fourteen simulated and four real-world robotic tasks demonstrate that the proposed approach achieves superior generalization performance in zero-shot scenarios.
📝 Abstract
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.