🤖 AI Summary
This study addresses the limited robustness of Vision-Language-Action (VLA) models in closed-loop control under action perturbations by proposing a dendrite-inspired spiking action architecture. The approach introduces dendritic spiking dynamics into VLA models, enabling modular feature processing through multi-branch sparse connectivity and heterogeneous decay factors. Furthermore, a neuron-level inhibitory gating mechanism is employed to filter noise interference while preserving historical task information. Evaluated on the LIBERO benchmark, the proposed method achieves a nominal success rate of 91.6% and maintains 87.35% under perturbations, corresponding to a performance retention rate of 95.4%. These results demonstrate significant improvements over existing mainstream baseline models.
📝 Abstract
Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring heterogeneous, learned decay factors. Furthermore, to suppress unreliable state updates while preserving task-relevant historical information, we introduce a neuron-wise inhibitory gate that adaptively regulates the admission of new multimodal evidence into dendritic states prior to somatic dynamics. We evaluate DS-VLA on all four LIBERO suites under both nominal rollouts and a unified closed-loop action-perturbation protocol. DS-VLA achieves a 91.6\% average nominal success rate and an 87.35\% average perturbed success rate, retaining 95.4\% of its nominal performance. Under the same reported perturbation setting, OpenVLA-OFT, FAST, $\pi_0$, and GR00T achieve 39.45\%, 23.90\%, 28.55\%, and 30.75\%, respectively. A controlled ablation isolates the contribution of neuron-wise shared inhibition, while analyses of neural dynamics and post-perturbation trajectories associate robust performance with selective evidence suppression and effective behavioral recovery. Together, these results demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for robust embodied intelligence beyond merely scaling vision-language backbones or generative action decoders.