The Role of Fine-grained Harm Signals in LLM Safety

📅 2026-09-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过移除通用有害性表示并使用激活导向方法,探讨了类别特异性成分在LLM安全性中的作用,发现其对理解模型安全至关重要。
📝 Abstract
Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs' downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification.
Problem

Research questions and friction points this paper is trying to address.

LLM Safety
Harm Signals
Category-specific Component
Innovation

Methods, ideas, or system contributions that make the work stand out.

category-specific component
general harmfulness representation
activation steering
internal alignment
🔎 Similar Papers
No similar papers found.