Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Families

📅 2026-06-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of emergent, poorly understood misalignments in fine-tuned language models that lead to unsafe code generation. The authors propose an interpretable activation-space framework that identifies actionable directions shared across architectures to detect and intervene on such misaligned behaviors. They demonstrate for the first time that internal misalignment directions exhibit causal specificity and manipulability, and uncover an asymmetric topological structure and a two-layer specificity mechanism underlying cross-architecture transfer. By integrating mean-difference directions, causal interventions, ridge regression mapping, and controlled safe-code experiments, they achieve 99.6% activation separation across four model families. Intra-model interventions reduce code leakage risk by 21–51 points, while cross-architecture transfer suppresses it by up to 46 points, albeit with limited specificity.
📝 Abstract
Fine-tuning language models on insecure code induces emergent misalignment with poorly understood internal structure. We investigate whether this misalignment corresponds to a causally actionable activation-space direction shared across architectures. Across four instruction-tuned model families (Qwen2.5-1.5B, Gemma-2-2B, Llama-3.2-1B, Ministral-3-3B) finetuned identically, a difference-in-means direction achieves 99.6% separation of aligned and misaligned activations at each model's final layer. Causal steering by subtracting this direction reduces code spillover by 21-51 points, while a secure-code control confirms content specificity. Cross-architecture transfer via ridge regression maps yields large behavioral suppression (up to 46 points) but fails specificity controls as random and orthogonal directions perform comparably. We identify a two-tier specificity structure: within-model directions are causally specific and actionable; cross-model directions are causally real but non-specific. An asymmetric transfer topology emerges, with Gemma and Qwen acting as geometric donors and Llama as a receiver. These findings define the limits of linear cross-architecture correction and recommend within-model probing for auditing.
Problem

Research questions and friction points this paper is trying to address.

emergent misalignment
activation directions
language model alignment
cross-architecture transfer
code spillover
Innovation

Methods, ideas, or system contributions that make the work stand out.

emergent misalignment
activation steering
cross-architecture transfer
causal specificity
linear intervention
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Abdul Rafay Syed
Department of Computer Science, Universität des Saarlandes, Saarbrücken, Germany