The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that safety instructions in large language models struggle to balance harmful and benign queries, often leading to excessive over-refusal. To this end, we propose a spectrum optimization-based safety regulation method. By modeling context as a multiplicative operator, our approach modulates its principal eigenvalue via eigendecomposition, enabling continuous control over safety influence intensity. Furthermore, we design a contrastive safety loss coupled with a suppression weighting mechanism to elucidate the intrinsic role of eigenvalues as a continuous safety knob. Integrating Transformer architectures with multi-objective optimization, the proposed method achieves Pareto improvements in both attack success rate and over-refusal rate, significantly enhancing the precision and efficiency of safety instruction regulation.
📝 Abstract
Context tokens in a transformer-based language model can be absorbed into the model's weights as a multiplicative operator. We study this operator in the setting of safety instructions and show that influencing its dominant eigenvalue modulates how strongly the instruction shapes generation. We derive a Contrastive Safety Loss with a suppression weight that controls the tradeoff between emphasizing the safety instruction on harmful queries while suppressing it on harmless queries. Varying the suppression weight maps a relationship between the attack success and the over-refusal rates, supporting the hypothesis that the operator's eigenvalue acts as a continuous dial for the instruction's influence. Moreover, this relationship holds relatively independently of how the Safety Loss is parameterized, yielding Pareto-improved safety instructions for appropriate values of suppression weight.
Problem

Research questions and friction points this paper is trying to address.

Safety Instructions
Over-refusal
Attack Success Rate
Large Language Models
Safety Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Operator
Spectral Optimization
Contrastive Safety Loss
Eigenvalue Modulation
Over-refusal
🔎 Similar Papers