Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of conventional safety auditing for large language models, which relies on behavioral evaluation, requires model execution, and is constrained by benchmark coverage. We propose TRACE, a weight-level auditing framework that reveals, for the first time, the geometric order of supervised fine-tuning (SFT) objectives within the parameter space. By constructing geometric references, conducting Spearman correlation analysis, and performing prototype matching in the weight space, TRACE enables leakage-free risk assessment without querying the model or accessing data. Experiments demonstrate that the proposed coordinates achieve correlation coefficients exceeding 0.986 with harmful objectives. The method remains stable across varying model scales, effectively correlates with attack success rates, and exhibits superior performance in low-ratio harmful fine-tuning scenarios.
📝 Abstract
Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at https://anonymous.4open.science/r/Code4TRACE-54D3.
Problem

Research questions and friction points this paper is trying to address.

Safety Auditing
Supervised Fine-Tuning
Checkpoint Updates
Large Language Models
Harmful Compliance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Checkpoint Auditing
Supervised Fine-Tuning
Weight-space Geometry
Safety Alignment
TRACE
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Z
Ziqun Bao
SEI, East China Normal University
X
Xinyu Zhang
SEI, East China Normal University
Y
Yuchen Shao
SEI, East China Normal University
Chengcheng Wan
Chengcheng Wan
East China Normal University
Software engineeringsystem optimizationmachine learning