🤖 AI Summary
This study addresses the high latency and energy consumption of existing Vision-Language-Action (VLA) models that rely on large Transformers, which hinder deployment on resource-constrained platforms. To this end, we propose the first spiking-driven VLA framework. It employs spiking visual and instruction encoders, introduces a multi-winner spike fusion mechanism with bidirectional Top-k routing to suppress background interference, and integrates spiking cross-attention with an action chunking Transformer to generate continuous control commands, enabling end-to-end training via sparse event-driven computation. Experiments demonstrate that our method achieves highly competitive performance on the LIBERO and Meta-World benchmarks while requiring fewer parameters and lower inference energy. This work lays a foundation for neuromorphic embodied intelligence.
📝 Abstract
Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.