Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Diffusion-based vision-language-action (VLA) models face significant computational overhead in embodied intelligence, hindering their deployment on edge devices due to stringent latency and power constraints. To address this challenge, this work proposes Deltoris, an algorithm-hardware co-designed inference framework that introduces a novel temporally aware bit-level sparsity algorithm and a speculative execution mechanism to drastically reduce redundant computation and amortize data loading costs. Complementing the algorithmic innovations, Deltoris features a customized one-dimensional systolic bit-serial processing element (PE) array that effectively mitigates workload imbalance. Experimental results demonstrate that Deltoris achieves up to 34.2ร— speedup over mobile GPUs and 6.1ร— improvement over state-of-the-art accelerators, while maintaining comparable inference accuracy and significantly enhancing energy efficiency.
๐Ÿ“ Abstract
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action
diffusion-based models
real-time inference
edge devices
latency constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

bit-level sparsity
speculative inference
algorithm-hardware co-design
diffusion-based VLA
temporal-aware acceleration