H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出H-VLA模型,通过高层关键动作推理与低层运动生成解耦来解决现有VLA模型在空间变化下的脆弱性问题。
📝 Abstract
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
semantic reasoning
spatial variations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Vision-Language-Action Model
Key-Action Reasoning
Motion Planning
Unified Camera-Centric Action Space
Two-Stage Training Strategy
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiongfeng Peng
Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB), China
Lu Xu
Lu Xu
Postdoc, Riken AIP
deep learningmachine learningcomputer vision
Yandong Wang
Yandong Wang
Citadel Securities
Big DataNoSQL storesMachine LearningHigh-Frequency Trading Systems
Jiaqian Yu
Jiaqian Yu
Samsung R&D Institute China - Beijing
Machine LearningComputer Vision
Z
Zirui Zheng
Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB), China
Y
Yamin Mao
Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB), China
Weiming Li
Weiming Li
Principal Engineer, Samsung Electronics
Computer VisionAugmented RealityComputational Imaging and Display
Inseop Chung
Inseop Chung
Samsung AI Center, DS Division, South Korea
H
Hyun-woong Cho
Samsung AI Center, DS Division, South Korea
J
Jaewook Yoo
Samsung AI Center, DS Division, South Korea
D
Dongwook Lee
Samsung AI Center, DS Division, South Korea
D
Daehyun Ji
Samsung AI Center, DS Division, South Korea
C
Chao Zhang
Advanced Research Lab, Samsung R&D Institute China-Beijing (SRCB), China