Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the safety risks inherent in Mixture-of-Experts (MoE) vision-language models and the over-refusal issues induced by conventional intervention strategies. We propose a lightweight, non-intrusive safety detector that departs from the traditional paradigm of manipulating internal model states. Instead, our method identifies unsafe requests by reading router logit signals that naturally emerge during the prefill phase, requiring no modifications to model parameters or expert routing mechanisms. Experimental results demonstrate that this approach significantly reduces error rates on the HoliSafe benchmark across Qwen3-VL and Kimi-VL models. Furthermore, it generalizes effectively to out-of-distribution benchmarks such as MISHard, achieving a favorable balance between safety alignment and model utility.
📝 Abstract
Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture-of-experts (MoE) VLMs become increasingly common, recent work has explored various safety interventions, including prompting, supervised fine-tuning, and routing-based expert steering. However, these methods show inconsistent improvements across models and evaluation distributions, and the intervention into model behavior or internal states introduce safety-utility tradeoffs by over-refusal. Rather than manipulating internal states to steer model behavior, we instead ask whether routing states can serve as diagnostic signals for multimodal safety. We find that router logits indeed provide highly predictive signals of whether a multimodal input is safe or not. Motivated by this observation, we introduce a lightweight router-logit safety detector that reads out routing signals during prompt prefill and identifies unsafe requests before generation, without modifying model parameters or expert routing. Across Qwen3-VL and Kimi-VL, the proposed detector substantially reduces safety errors on the HoliSafe benchmark and resoundingly generalizes to out-of-distribution safety benchmarks featuring different safety patterns, including MISHard and MM-SafetyBench. The success of the proposed router-logit detector also suggests a broader perspective on model internals: rather than focusing only on manipulating internal components to steer behavior, simply reading naturally emerging signals and linking them to an external safety mechanism can provide a simple, effective, and non-intrusive complement to existing safety interventions.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Mixture-of-Experts
Multimodal Safety
Safety-Utility Tradeoff
Over-refusal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Router Logits
Mixture-of-Experts
Multimodal Safety
Vision-Language Models
Lightweight Detector
🔎 Similar Papers
No similar papers found.