🤖 AI Summary
This study addresses the limitation of existing LLM interpretability tools that focus exclusively on activated features while overlooking safety-relevant components within inactive neurons, thereby undermining jailbreak detection. We propose the Counterfactual Activation Potential (CAP) metric and the CSFD algorithm, leveraging mechanistic interpretability and transcoder analysis to mine suppressed safety features at scale. This work is the first to quantify the safety contributions of inactive features, revealing that the core mechanism underlying jailbreak attacks lies in suppressing safety features rather than merely activating harmful ones. Multi-model experiments demonstrate that ablating candidate features significantly increases violation rates, while under natural jailbreaks, high-CAP features exhibit intensified suppression and substantial activation declines. These findings confirm that our approach effectively bridges critical blind spots in current interpretability frameworks.
📝 Abstract
Mechanistic interpretability has emerged as the primary means to understand safety behavior of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it). To find suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. A significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. Under natural jailbreaks, the suppression acting on the highest-CAP features rises 2-4x, and their activation correspondingly falls by up to 80%. Amplifying a feature's suppressors pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experiments span five Gemma, Qwen, and Llama models across various parameter sizes. Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.