DNAlign: Dynamic Null-Space Safe Alignment for LLMs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in safety alignment for large language models (LLMs), where high computational costs often compromise core knowledge and degrade utility. To overcome this, we propose a lightweight dynamic null-space safety alignment framework grounded in control theory. By conceptualizing LLMs as dynamical systems, our approach leverages value function training and null-space projection techniques to strictly constrain alignment perturbations within harmful subspaces, thereby enabling precise interventions while preserving general capabilities. As a pioneering effort, this work introduces a dynamic null-space mechanism that simultaneously safeguards generation diversity and model security. Experimental results demonstrate that the proposed method significantly reduces harmful outputs while effectively maintaining factual accuracy and linguistic fluency. Overall, it outperforms existing baselines without sacrificing generative diversity.
📝 Abstract
Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a persistent trade-off between safety and utility. We propose DNAlign, a lightweight alignment framework that integrates control-theoretic optimization with null-space projection. By treating the LLM as a dynamic system, the proposed framework introduces controllable perturbations to steer generation toward safe behavior. A key component is the projection module, which restricts these perturbations to the harmful-related subspace derived from neutral hidden states, thereby preserving general knowledge and response quality. A value function trained on human preference data adaptively optimizes the control signals to align with human safety preferences. Extensive evaluations across multiple LLM backbones demonstrate that our framework consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility. It achieves superior overall performance compared to prior alignment baselines without sacrificing generation diversity. These results indicate that the proposed framework provides an effective and practically deployable solution for safe LLM alignment. Code is available at https://anonymous.4open.science/r/DNAlign.
Problem

Research questions and friction points this paper is trying to address.

Safety Alignment
Large Language Models
Safety-Utility Trade-off
Reliable Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Null-Space Projection
Control-Theoretic Optimization
Safety Alignment
Lightweight Framework
Value Function
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jisheng Dang
School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China
Y
Yushuo Zhao
School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China
D
Dewei Liu
School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China
Junfeng Fang
Junfeng Fang
National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science
B
Bimei Wang
School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China
T
Tiantian Rao
School of Information Science and Engineering, Lanzhou University, Lanzhou 730000, China
Hong Peng
Hong Peng
Vice Professor of Physice, Lanzhou University
EEGAffective ComputingDepressionAnxiety neurosis
B
Bin Hu
School of Medical Technology, Beijing Institute of Technology, Beijing 100081, China
T
Tat-Seng Chua
National University of Singapore, Singapore 119077