๐ค AI Summary
Diffusion-based vision-language-action (VLA) models face significant computational overhead in embodied intelligence, hindering their deployment on edge devices due to stringent latency and power constraints. To address this challenge, this work proposes Deltoris, an algorithm-hardware co-designed inference framework that introduces a novel temporally aware bit-level sparsity algorithm and a speculative execution mechanism to drastically reduce redundant computation and amortize data loading costs. Complementing the algorithmic innovations, Deltoris features a customized one-dimensional systolic bit-serial processing element (PE) array that effectively mitigates workload imbalance. Experimental results demonstrate that Deltoris achieves up to 34.2ร speedup over mobile GPUs and 6.1ร improvement over state-of-the-art accelerators, while maintaining comparable inference accuracy and significantly enhancing energy efficiency.
๐ Abstract
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices.
In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.