🤖 AI Summary
This study addresses the physical safety risks posed by visual input failures, specifically blackouts and frame freezing, to Vision-Language-Action (VLA) models, which can trigger unexpected interactions such as extreme joint behaviors. Through systematic VLA model testing, selective intervention experiments, embedding replacement, and real-world robotic trials, this work analyzes differentiated failure modes and evaluates the efficacy of proprioceptive compensation mechanisms alongside two mitigation strategies. The findings reveal the inherent limitations of proprioception under visual deprivation and propose a robust design paradigm leveraging residual information. Furthermore, this research demonstrates that while the proposed mitigation measures improve task success rates, they concurrently introduce increased interference. Ultimately, these insights provide empirical foundations for designing safety strategies in embodied intelligence systems.
📝 Abstract
Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $π0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even when task-success rates are similarly low: freezing causes more extreme joint behavior, whereas blackout after gripper closure can cause more object drops, most markedly without proprioception. Selective intervention studies reveal that proprioception (current robot state) partly compensates for the removed robot depictions and reduces non-target contact. However, it cannot sufficiently restore task success when wrist-view object information is removed, even when aided by the remaining scene view. We then evaluate two mitigation approaches: camera-blackout training and training-free replacement of faulty visual embeddings. Both improve task success in selected conditions, but can increase unintended contact or disturbance to surrounding objects. Real-robot trials further show that successful execution under camera faults can still involve unintended physical interactions. These findings motivate designing VLA policies that use the robot and object information still available under camera faults to limit hazardous motion.