TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment

📅 2026-06-26
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of safety alignment caused by downstream task adaptation during large model fine-tuning services, where existing remediation approaches often compromise task performance. To overcome this limitation, this work proposes an offline safety patch learning framework that transforms conventional online calibration into offline optimization by generating patches through simulated harmful fine-tuning trajectories. Notably, training a single universal patch suffices to accommodate diverse user-specific fine-tuned checkpoints, enabling interaction-free safety restoration. Evaluations across six benchmarks demonstrate that the proposed method significantly dominates the safety-utility Pareto frontier, improving the safety rate by 77 percentage points while preserving equivalent task utility. These results effectively resolve the fundamental challenge of simultaneously maintaining safety alignment and downstream performance in fine-tuning scenarios.
📝 Abstract
Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.
Problem

Research questions and friction points this paper is trying to address.

Safety Alignment
Supervised Fine-Tuning
Safety-Utility Trade-off
Safety Patch
Post-Training Realignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Safety Patch Learning
Offline Optimization
Trajectory Simulation
Post-Training Realignment
Safety-Utility Trade-off
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.