🤖 AI Summary
This study addresses the safety vulnerabilities of Vision-Language-Action (VLA) models under single-step observation perturbations by focusing, for the first time, on instantaneous perturbation threats. We propose CARE, a dynamic execution length selection method grounded in predictive consistency. By integrating a consistency-checking algorithm with an optimized action chunk execution strategy, CARE dynamically adjusts the action execution horizon based on prediction outcomes to enhance robustness. Experimental results demonstrate that CARE significantly reduces computational overhead while effectively defending against transient observation perturbations. Furthermore, it maintains superior performance in perturbation-free scenarios. This work establishes a novel paradigm for the reliable deployment of VLA models in real-world environments.
📝 Abstract
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which observations are severely perturbed only briefly within an episode, remains an underexplored safety threat. To address this gap, this work investigates robustness to one-step perturbations applied at a single time step per episode. Our experiments reveal that these perturbations substantially degrade VLA performance and that their impact depends on the action chunk execution length. Based on them, we propose CARE, which dynamically selects the execution length based on consistency with the previously predicted action chunk. CARE improves robustness with low computational overhead while preserving clean performance by selecting shorter execution lengths only under perturbations.