🤖 AI Summary
This study addresses the high inference latency of Vision-Language-Action (VLA) policies caused by iterative denoising, which hinders real-time control. We propose an urgency-aware denoising framework that exploits the asymmetry between action consumption and generation order to dynamically allocate computational resources based on execution urgency, prioritizing the release of critical actions while concurrently refining subsequent ones. Furthermore, trajectory harmonization and ghost action correction mechanisms are introduced, combined with flow matching and dynamic error compensation techniques to ensure action consistency and precision. Experimental results demonstrate that the proposed method achieves a 1.89× average speedup in action availability latency, significantly outperforming existing state-of-the-art baselines while maintaining comparable task success rates.
📝 Abstract
Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies. We exploit this asymmetry to introduce Urgency-Aware Denoising (UAD), a novel inference-time framework that allocates denoising computation according to when each action is physically needed. UAD releases time-critical urgent actions after fewer denoising steps while overlapping the continued background refinement of tail actions with physical execution. However, heterogeneous denoising introduces two key challenges: early-release errors in urgent actions and trajectory inconsistency in tail actions. UAD elegantly resolves both through two core mechanisms: Trajectory Reconciliation, which reconstructs unified internal state evolution to restore joint denoising coherence without additional model evaluations, and Ghost Action Correction, which leverages non-executed ghost continuations to dynamically compensate for early-release errors across remaining executable actions. Extensive evaluations across multiple VLA architectures, simulation benchmarks, and real-world manipulation tasks demonstrate that UAD achieves up to a 1.89x speedup in average action availability latency while maintaining comparable success rates to vanilla inference with optimal denoising budget, offering a more favorable success-latency trade-off than state-of-the-art VLA acceleration baselines.