Multimodal Safety Evaluation Should Measure Controllability Beyond Classification
This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.