See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the cross-domain safety degradation, termed emergent misalignment, induced by domain adaptation in large language models, where existing static analyses and defenses struggle to balance safety with utility. To overcome this limitation, we propose a dynamic second-order geometric diagnostic mechanism based on Grassmannian manifold projection that traces training trajectories and reveals the concentration of harmful gradients on semantic pivot tokens. Building upon these insights, we construct a parameter-level orthogonal projection framework that precisely eliminates the harmful gradient subspace, effectively decoupling safety from utility. This work establishes the first dynamic geometric mitigation paradigm for such challenges. Experiments demonstrate that our approach suppresses free-form generation misalignment by up to 80% on Qwen2.5-14B, while further validating its effectiveness in controlling behavioral safety hallucinations across multiple open-source models.
📝 Abstract
Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens. Grassmannian projections show that, in most settings, harmful-safe gap widens mainly because safe-gradient overlap declines. Leveraging these insights, we introduce a parameter-level Geometric Mitigation Framework that orthogonally projects empirical harmful gradient subspace out of parameter updates. On Qwen2.5-14B-IT, our defense suppresses free-generation EM by up to 80.0%; across the other three of four open-weight instruction-based model families (3B--20B), where single-layer behavioral EM is already near zero, teacher-forced evaluation shows same harmful subspace controls the conditional support of frozen EM responses. Crucially, these diagnostics unmask the illusion of behavioral safety: the same subspace remains measurable and steerable in models where behavioral EM is near zero. Code: https://github.com/WeiqiaoQUE/mechanistic-emergent-misalignment.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emergent Misalignment
Second-order Geometric Analysis
Grassmannian Projections
Geometric Mitigation Framework
Orthogonal Projection
W
Weiqiao Que
School of Computer Science and Technology, East China Normal University, China
R
Ruizhe Li
School of Computer Science, University of Birmingham, UK
Chengyu Wang
Chengyu Wang
Alibaba Group
Natural Language ProcessingLarge Language ModelMulti-modal Learning
D
Dakan Wang
Exacity Inc.
Emine Yilmaz
Emine Yilmaz
University College London
Information RetrievalNatural Language ProcessingMachine Learning
X
Xiaofeng He
School of Computer Science and Technology, East China Normal University, China