🤖 AI Summary
This study investigates whether the safety scores of frozen vision-language models (VLMs) reflect genuine hazard perception or merely respond to scene features. Leveraging a frozen CLIP model and reinforcement learning trajectory generation, we design controlled experiments that isolate contact onset and systematically dissect the mechanisms underlying marginal safety scores through multivariate manipulations of encoders, viewpoints, and captions. Our findings reveal that score degradation stems from shifts in pre-collision scene similarity rather than hazard recognition, demonstrating that such scores fundamentally capture text-scene matching instead of authentic danger awareness. Consequently, this work establishes that improvements in policy performance do not necessarily indicate genuine safety understanding in these models, offering critical mechanistic insights for VLM safety evaluation.
📝 Abstract
Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.