Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.
📝 Abstract
VLM safety is commonly evaluated through input- and output-level classification. Such classification is necessary, but it does not reveal whether a safety state is accessible or controllable inside the model. We argue that multimodal safety evaluation should therefore report a \emph{controllability profile} alongside behavioral classification, separating representation-level detectability, cross-modal specificity, intervention sensitivity, and benign-preserving selectivity. Using implicit toxicity as a stress case, we instantiate this profile on LlavaGuard and Qwen3.5 with sparse feature decompositions. LlavaGuard admits localized handles with a narrow benign-preserving intervention range and modest downstream safety gains, whereas Qwen3.5 supports strong representation-level readout but no comparable selective-control regime under the tested operators. These results show that internal readout and controllability can diverge. Future multimodal safety benchmarks should therefore report not only behavioral safety metrics, but also whether safety-relevant internal signals can be intervention-tested and controlled within a validated operating range.
Problem

Research questions and friction points this paper is trying to address.

multimodal safety evaluation
controllability
vision-language models
behavioral classification
internal representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Safety Evaluation
Controllability Profile
Sparse Feature Decomposition
Implicit Toxicity
Intervention Sensitivity
🔎 Similar Papers
No similar papers found.