🤖 AI Summary
This study addresses the vulnerability of safety refusal mechanisms in embodied Vision-Language Model (VLM) planners to minor environmental variations, highlighting their susceptibility to circumvention by non-adversarial perturbations. To investigate this, the work introduces and quantifies the novel concept of task "plasticity," probing safety boundaries through everyday object perturbations. Furthermore, it constructs a predictive model leveraging signal composition analysis of open-source VLMs to conduct a large-scale benchmark evaluation of task safety. The findings reveal that non-adversarial environmental changes can undermine safety alignment, with experiments demonstrating that 20.2% of refused tasks can be reversed. Notably, the proposed predictive model achieves a pre-deployment risk assessment accuracy 2.4 times higher than a random baseline, offering a practical approach for evaluating the robustness of embodied VLM planners prior to deployment.
📝 Abstract
Embodied vision-language models (VLMs) are increasingly deployed as high-level planners for robots because they generalize across diverse environments. However, this requires their safety alignment to also hold in unseen environments. Existing red-teaming assumes an adversary who optimizes the prompt, the pixels, or text in the environment, and existing benchmarks ask whether a planner recognizes or mitigates a hazard in a fixed scene. Neither asks whether a refusal the planner has already given survives an ordinary change to the environment. We ask that question by placing a single everyday object into the environment, with no pixel, gradient, or prompt under adversarial control. On $846$ tasks that a constitution-guarded planner initially refuses, we find $20.2\%$ of tasks can be flipped to compliance by one or more objects, and the number of objects differs from one task to another. In addition, the object need not be chosen for the task, i.e., items drawn from a fixed list, with no knowledge of the environment or the instruction, bypass safety about as often as items proposed for the specific task. We qualitatively contrast the tasks bypassed most and least often and find that the distinction lies in how conspicuous the hazard is in the instruction and environment. Susceptibility to safety bypass is therefore a property of the task, which we call its \emph{malleability}, and we show that it can be predicted before the target is ever queried. A composite of signals read from a small open-source VLM identifies malleable tasks $2.4\times$ as often as picking at random. Everyday objects, whether placed by an adversary or introduced by ordinary rearrangement of the environment, are thus sufficient to overturn a refusal. Because susceptibility is determined by how a task is specified, we recommend assessing malleability per task prior to deployment.