SafeEvo: Deciphering the Safety Alignment Mechanism and Evolution in Language Models

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight of internal mechanism evolution in pretrained models during safety alignment by proposing a novel interpretability framework that analyzes safety mechanisms through the lens of sparse subgraph circuits. Methodologically, it employs optimization algorithms to extract and trace the evolution of refusal circuits, introducing Safety Circuit Alignment—a technique that strictly constrains parameter updates within these circuits to achieve precise local alignment. The analysis reveals that the "alignment tax" stems from global updates interfering with utility parameters. Experimental results demonstrate that, compared to conventional alignment methods, the proposed approach reduces harmfulness by 63.21% and false refusal rates by 58.44%, while preserving 99.58% of original capabilities, thereby significantly enhancing the overall trade-off between safety and utility.
📝 Abstract
Safety interpretability advances the study of Large Language Model (LLM) alignment from behavioral constraints driven by data or algorithms towards a deeper understanding of internal mechanisms. However, existing works have focused primarily on safety-related representations, attention heads, or neurons after alignment, while largely overlooking the safety mechanisms in pretrained-only models and their evolution across alignment checkpoints. To address this, we propose SafeEvo, an interpretability framework from the circuit (sparse subgraphs of an LLM) perspective. SafeEvo first applies an optimization-based extraction algorithm to identify weak refusal circuits in pretrained base LLMs that can independently express refusal behavior. Causally ablating these circuits completely eliminates the base model's refusal of harmful inputs. SafeEvo then traces the evolution of refusal circuits across successive alignment checkpoints and finds that their structures change progressively, suggesting that the alignment tax may result from refusal-circuit updates affecting utility-related parameters. To validate this, SafeEvo introduces Safety Circuit Alignment (SCA), which confines safety updates to the refusal circuits. Experiments across three LLMs and two alignment algorithms show that, on average, SCA outperforms vanilla alignment in three aspects: \textbf{(1) stronger alignment}, lowering harmfulness score by 63.21\%; \textbf{(2) less over-refusal}, yielding a 58.44\% decrease in refusal rates for benign queries; and \textbf{(3) better utility}, retaining 99.58\% of the original model capabilities.
Problem

Research questions and friction points this paper is trying to address.

safety alignment
interpretability
refusal circuits
language models
alignment tax
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Interpretability
Circuit Extraction
Refusal Circuits
Safety Circuit Alignment
Alignment Tax
🔎 Similar Papers
No similar papers found.