Fast Plans, Faithful Actions: Closing the Planning-Execution Gap in Hierarchical Vision-Language-Action Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high planning latency and insufficient plan utilization in hierarchical Vision-Language-Action (VLA) models. To this end, it proposes waypoint-aligned block-autoregressive decoding and normalized goal modulation techniques. By incorporating inter-layer path constraints, phase-gating mechanisms, and anti-shortcut training strategies, the method effectively optimizes the alignment between the planner and the executor, substantially reducing inference overhead while enhancing plan utilization. Experimental results demonstrate that the proposed approach reduces planning latency by 8.7× and achieves a 98.45% success rate on the LIBERO benchmark, while also exhibiting strong performance in bimanual robotic tasks. Overall, this work provides an efficient solution for improving both the real-time responsiveness and effectiveness of VLA systems.
📝 Abstract
Hierarchical vision-language-action (VLA) systems consist of a high-level vision-language planner and a low-level action expert that generates continuous actions. This hierarchical design has practical value only if the planner can generate plans fast enough to meet real-time control requirements, and the resulting plans actually contribute to the generation of action. We study one such system, a waypoint hierarchy pipeline adapted from $π_{0.5}$, and find that neither requirement is satisfied. This baseline relies on token-level autoregressive decoding (Token-AR) to generate a waypoint plan, requiring 57 very expensive vision-language model (VLM) forward passes. However, we find that erasing the waypoint endpoints has little effect on task success. Two findings reveal the misalignment of planner-executor: the planner generates outputs at an excessively fine granularity, and the executor underuses plans as a control condition. We address the latency issue with waypoint-aligned block-autoregressive decoding (Block-AR), and plan underuse issue with normalized goal modulation (NGM), a layer-wise goal path constrained by phase gating and anti-shortcut training so that the waypoint influences action generation maintaining other signals. Our method reduces the maximum number of VLM forward passes from 57 to 8 on LIBERO, including one prefix prefill, and achieves an $8.7\times$ reduction in planning latency on a Rokae dual-arm robot. With normalized goal modulation and anti-shortcut training, Block-AR's success rate on LIBERO-Long increases from 91.0% to 96.2%, while its average success rate across the four suites increases from 95.85% to 98.45%. On three bimanual tasks with this robot, success rates remain comparable across methods.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Models
Hierarchical Planning
Planning-Execution Gap
Real-time Control
Plan Underuse
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Vision-Language-Action Models
Block-Autoregressive Decoding
Normalized Goal Modulation
Anti-Shortcut Training
Planning-Execution Gap
🔎 Similar Papers
No similar papers found.