🤖 AI Summary
This study addresses the imbalance between inference efficiency and accuracy in Vision-Language-Action (VLA) models, caused by the dynamic drift of conditional attention during diffusion denoising. To this end, we propose a stage-aware hierarchical action generation framework that reveals how conditional focus shifts throughout the denoising process. By leveraging partially denoised actions to construct hierarchical interfaces, the framework decouples long-horizon planning from high-frequency local corrections. Furthermore, it integrates flow matching with a lightweight action refiner to achieve computation amortization and closed-loop correction. Evaluated on the LIBERO benchmark, our method improves the success rate to 97.8% while reducing inference latency to 44.2 ms, significantly lowering computational costs without compromising performance.
📝 Abstract
Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.