ASCENT: First-Order Optimal Fine-Tuning with Recalibration for Safety--Utility Co-Enhancement

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of safety alignment and the inherent difficulty of jointly improving safety and utility during large language model fine-tuning. To this end, we propose ASCENT, a framework that achieves dual enhancement by alternating between first-order optimal safety-aware calibration and task optimization. Theoretically, we prove that updates based on gradient singular components maximize safety variation and derive a unique safety-preserving task update direction to overcome static subspace limitations. Algorithmically, we design a periodic calibration strategy integrating singular value decomposition with Frobenius norm constraints. Experimental results demonstrate that ASCENT improves downstream utility by 20.3% while reducing the attack success rate by 35.5%, achieving state-of-the-art performance.
📝 Abstract
Supervised fine-tuning can substantially improve the downstream utility of large language models (LLMs) but may compromise their safety. Existing safety-preserving methods constrain downstream updates using safety-related parameters or subspaces, but mainly focus on safety preservation rather than joint safety and utility enhancement, lack a theoretical characterization of the optimal safety-related subspace and safety-preserving task update, and typically rely on a static safety subspace that may become outdated during fine-tuning. To address these limitations, we propose ASCENT, a downstream fine-tuning framework for safety--utility co-enhancement through first-order optimal safety-aware periodic calibration and task optimization. We model safety as a function of LLM parameters $S(θ)$ and use its first-order approximation to characterize safety changes under parameter updates. Under a fixed rank and Frobenius-norm budget, we prove that the update constructed from the top-$r$ singular components of the safety-function gradient maximizes the estimated safety change, and use it for periodic calibration to preserve and improve safety. We further derive a unique safety-preserving task update that stays close to the original task update while penalizing negative effects on the estimated safety change. ASCENT alternates these optimal task and calibration updates to jointly enhance safety and utility. Experiments across multiple LLM families and downstream tasks show that ASCENT improves downstream utility by up to 20.3\% and reduces attack success rate by up to 35.5\%, achieving state-of-the-art safety and utility across all evaluated settings. Our code is available at https://github.com/ZJU-LLM-Safety/ASCENT.
Problem

Research questions and friction points this paper is trying to address.

Supervised fine-tuning
Safety-utility co-enhancement
Large language models
Safety preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety-Utility Co-Enhancement
First-Order Approximation
Periodic Calibration
Supervised Fine-Tuning
Singular Value Decomposition
🔎 Similar Papers