🤖 AI Summary
This study addresses the vulnerability of existing safety-aligned defenses in vision-language models to intent-preserving input reconstruction attacks. We propose SteerProbe, a black-box attack framework that leverages calibration learning to perform activation-steered local reconstruction of multimodal inputs. By establishing sufficient conditions for traversing safety boundaries, SteerProbe enables efficient, automated output-side attacks. Experimental results demonstrate that our method significantly increases the average harmfulness rate from 7.43% to 18.36%, exposing critical robustness deficiencies in current defense mechanisms. This work provides a novel perspective for evaluating and enhancing the safety of multimodal large language models.
📝 Abstract
Activation steering offers an inference-time defense for vision--language models (VLMs) by modifying intermediate representations without updating backbone parameters. However, protection on benchmark inputs may not persist across alternative expressions of the same harmful request. We investigate this gap using fixed textual, visual, and joint reformulations designed to preserve the underlying intent, and find that these changes can bypass representative steering defenses. A complementary local analysis provides a sufficient condition under which a reformulation can cross a surrogate safety margin despite any admissible change in the local steering correction. We then introduce SteerProbe, an output-only black-box attack that learns to select effective reformulations for unseen requests from a shared calibration budget. Across three VLM backbones, two benchmarks, and three steering defenses, SteerProbe increases Harmful Rate in all defended settings using 500 total calibration queries per endpoint and benchmark, raising the average from 7.43% to 18.36%. These findings highlight that robustness on original benchmark inputs is insufficient to characterize the safety of steering defenses and motivate reformulation robustness as an important evaluation dimension. They further motivate steering mechanisms that preserve safety across intent-preserving multimodal variations while maintaining benign utility.