🤖 AI Summary
This work addresses the high inference latency of Vision-Language-Action (VLA) models on GPUs, which hinders real-time deployment, and the inadequacy of existing acceleration methods in fully eliminating weight and computational redundancy. To overcome these limitations, the paper proposes VQVLA, a novel framework that introduces motion-aware dynamic quantization (MotionVQ) and codebook-index-based vectorized GEMM. By leveraging spatiotemporal centroid reuse and dynamic precision selection within an algorithm-hardware co-design paradigm, VQVLA substantially reduces memory traffic and redundant computation. Experimental results demonstrate that VQVLA achieves a 6.5× speedup over the A100 GPU and a 2.8× speedup over Dadu-Corki while maintaining task success rates and incurring negligible accuracy loss.
📝 Abstract
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited.
In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.