Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM

๐Ÿ“… 2026-10-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limitation of existing sparse Mixture-of-Experts (MoE) models, which typically rely on activation frequency to identify safety-sensitive experts. Since frequency does not equate to influence, such approaches often yield imprecise localization. To overcome this, we propose a router gradient sensitivity metric for expert identification. By computing the gradients of sequence loss with respect to gating weights, our method precisely pinpoints critical safety experts and suppresses them to degrade model safety behavior without requiring retraining. We validate the effectiveness of this approach across multiple architectures, including OLMoE. Experimental results demonstrate that refusal rates are substantially reducedโ€”for instance, from 34% to 9%โ€”while generation quality remains uncompromised. This work establishes a novel paradigm for efficiently executing jailbreak attacks against MoE-based large language models.
๐Ÿ“ Abstract
Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the two budgets, router-gradient selection reduces refusals more than activation in 24 of 25 conditions, and more than a ten-trial random mean in all 25. The largest effect is in OLMoE, where refusals fall from 34 to 9 of 100 prompts (73.53% relative) with no degraded outputs, indicating substantive compliance rather than broken generation. After matching expert counts in every layer, gradient selection still produces greater refusal reduction than activation in 23 of 25 conditions, with two ties. An exploratory cross-model analysis links larger malicious-versus-benign concentration gaps to greater peak gradient effects (rho = 0.90; exact two-sided p = 0.083, n = 5). Together, the results support gradient selection under the tested budgets.
Problem

Research questions and friction points this paper is trying to address.

Sparse MoE
Safety Alignment
Expert Suppression
Activation Frequency
Router-Gradient Sensitivity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse MoE LLM
Router-Gradient Sensitivity
Safety-Sensitive Experts
Activation Frequency
Model Safety