🤖 AI Summary
This study addresses the high inference latency and deployment challenges of existing spiking vision-language-action (VLA) models, which typically require numerous time steps. To overcome these limitations, this work proposes an asynchronous spiking VLA framework that enables low-latency inference through an efficient artificial neural network to spiking neural network (ANN-SNN) conversion. The method introduces Dendritic Integrate-and-Fire (DIF) neurons coupled with an adaptive somatic firing mechanism to mitigate activation outliers, alongside a cross-component asynchronous execution strategy that overlaps computation to reduce synchronization overhead. Experimental results demonstrate that the proposed approach improves task success rates by 11.9% and optimizes path lengths by 12.6%, while achieving an 11.2-fold reduction in first-action latency. These findings establish a unified paradigm that simultaneously delivers high performance and computational efficiency for real-time robotic deployment.
📝 Abstract
ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9\% and 12.6\%, respectively, while reducing first-action latency by 11.2$\times$. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.