HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the disconnection among visual grounding, path planning, and execution decision-making in zero-shot robotic manipulation by proposing a keypoint-enhanced hierarchical architecture. Utilizing semantic 3D representations as a bridge, this method unifies visual grounding and language-conditioned planning through point-level interfaces. Specifically, it employs PointVLM for visual localization and integrates depth information to construct 3D representations, leverages 3DLLM for waypoint planning, and introduces a hybrid grasping module to achieve end-to-end control, all without requiring task-specific demonstration training. Experimental evaluations across fourteen simulated and four real-world robotic tasks demonstrate that the proposed approach achieves superior generalization performance in zero-shot scenarios.
📝 Abstract
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-relevant objects with image coordinates using a mixture of point annotations, segmentation-derived samples, robot observations, and visual question answering data. Depth measurements lift these predictions into a semantic 3D representation. A language planner, 3DLLM, uses this representation to specify end-effector waypoints and gripper commands, while a hybrid grasping module resolves local grasp poses. The evaluation covers 14 simulated manipulation tasks and four physical-robot tasks, together with ablations of the visual training data, spatial inputs, and grasp selection. Here, zero-shot execution refers to deployment without task-specific demonstration training; the visual model uses existing robot data during fine-tuning. This paper describes the original point-based formulation of the framework; its relationship to the subsequent GeneralVLA extension is detailed in the introduction.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot 3D Path Planning
Robot Demonstration Generation
Visual Grounding
Motion Planning
Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot 3D Path Planning
Hierarchical World Knowledge Model
Visual Keypoint Enhancement
PointVLM
3DLLM
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Guoqing Ma
Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Future Technology, University of Chinese Academy of Sciences; State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, CAS
Mingqi Yuan
Mingqi Yuan
PhD candidate at HKPU
Machine Learning
C
Chen Gao
Department of Electronic Engineering, Tsinghua University, China
J
Jiayu Chen
The University of Hong Kong, Hong Kong SAR, China
Shan Yu
Shan Yu
Institute of Automation, Chinese Academy of Sciences
Neuroscience