🤖 AI Summary
This study addresses the degradation of safety alignment caused by downstream task adaptation during large model fine-tuning services, where existing remediation approaches often compromise task performance. To overcome this limitation, this work proposes an offline safety patch learning framework that transforms conventional online calibration into offline optimization by generating patches through simulated harmful fine-tuning trajectories. Notably, training a single universal patch suffices to accommodate diverse user-specific fine-tuned checkpoints, enabling interaction-free safety restoration. Evaluations across six benchmarks demonstrate that the proposed method significantly dominates the safety-utility Pareto frontier, improving the safety rate by 77 percentage points while preserving equivalent task utility. These results effectively resolve the fundamental challenge of simultaneously maintaining safety alignment and downstream performance in fine-tuning scenarios.
📝 Abstract
Fine-Tuning-as-a-Service (FTaaS) platforms let users perform supervised fine-tuning (SFT) on customized data, but this pipeline can erode model safety alignment. To recover safety without re-running full alignment, existing realignment methods focus on calibrating the integration of safety patches into fine-tuned models. These methods exhibit a persistent safety-utility trade-off: weak repair leaves harmful behavior intact, while stronger repair increasingly damages the benign task. This paper shifts the focus from online calibration to offline patch learning and aims to learn a safety patch that restores safety while preserving task-specific capabilities. To this end, we propose TRACE, which simulates harmful SFT trajectories to produce progressively corrupted model states, and optimizes a safety patch simultaneously across these states. TRACE trains the safety patch during the offline stage, and reuses it across all user checkpoints without per-user calibration. We evaluate two representative models using three harmful SFT datasets, together with three utility benchmarks. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE improves the safety rate by up to 77 percentage points over the best baselines, while maintaining comparable utility to the undefended fine-tuned model.