🤖 AI Summary
This study addresses the vulnerability of Vision-Language-Action (VLA) models to adversarial patches, which can induce control failures through mechanisms that remain poorly understood. To elucidate these internal dynamics, this work leverages sparse autoencoders to interpret VLA representations and precisely identify attack-related features. Building upon this analysis, we propose a fine-tuning-free defense strategy that selectively suppresses the identified adversarial features only upon attack detection, thereby avoiding the degradation of nominal policy performance caused by continuous intervention. Experimental evaluations on the LIBERO-10 benchmark demonstrate that the proposed method significantly improves task success rates under intermittent adversarial attacks. By activating defensive measures solely when necessary, our approach effectively balances robustness against adversarial perturbations with the preservation of original model capabilities.
📝 Abstract
Adversarial patches can disrupt Vision-Language-Action (VLA) models by manipulating visual observations, leading to failures in robot control. However, it remains poorly understood which internal mechanisms underlie these failures and how targeted interventions can mitigate them. In this work, we mechanistically analyze VLA representations using a sparse autoencoder (SAE) and identify a feature whose activation strongly correlates with the presence of an adversarial patch. Based on this analysis, we suppress the identified feature at inference time only when a linear probe detects an attack. This intervention improves robustness without the cost of fine-tuning the VLA. We evaluate our method against VLA adversarial patch attacks on LIBERO-10. Conditional intervention improves success rate under intermittent attacks, whereas continuously applying the same intervention substantially degrades policy performance. These results show that attack-related internal representations can provide useful targets for VLA adversarial defense and that controlling when to intervene is important for limiting disruption to nominal policy behavior.