Spike-driven Vision-Language-Action Model

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high latency and energy consumption of existing Vision-Language-Action (VLA) models that rely on large Transformers, which hinder deployment on resource-constrained platforms. To this end, we propose the first spiking-driven VLA framework. It employs spiking visual and instruction encoders, introduces a multi-winner spike fusion mechanism with bidirectional Top-k routing to suppress background interference, and integrates spiking cross-attention with an action chunking Transformer to generate continuous control commands, enabling end-to-end training via sparse event-driven computation. Experiments demonstrate that our method achieves highly competitive performance on the LIBERO and Meta-World benchmarks while requiring fewer parameters and lower inference energy. This work lays a foundation for neuromorphic embodied intelligence.
📝 Abstract
Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action Model
Embodied Intelligence
Energy Efficiency
Resource-constrained Platforms
Spiking Neural Networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Neural Networks
Vision-Language-Action Model
Multi-Winner Spike Fusion
Spike Action Chunking Transformer
Embodied Intelligence
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Shuai Wang
Shuai Wang
University of Electronic Science and Technology of China
Spiking Neural Networks
M
Malu Zhang
University of Electronic Science and Technology of China
Mingquan Liu
Mingquan Liu
Hunan University
AIDD
Weihui Dai
Weihui Dai
University of Electronic Science and Technology of China
Dehao Zhang
Dehao Zhang
University of Electronic Science and Technology of China
Spiking Neural Network
J
Jieyuan Zhang
University of Electronic Science and Technology of China
Yimeng Shan
Yimeng Shan
Liaoning technical university
Spiking Neural NetworksNeuromorphic VisionSingle Object TrackingEvent Camera
Z
Zijian Zhou
University of Electronic Science and Technology of China
Y
Yang Yang
University of Electronic Science and Technology of China