A Mechanistic Investigation of Supervised Fine Tuning

📅 2026-05-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Although supervised fine-tuning (SFT) exerts only a subtle effect on the cosine similarity of hidden activations in large language models, their internal representations may nonetheless undergo substantial changes. This work proposes a high-resolution mechanistic analysis framework based on pretrained sparse autoencoders (SAEs), integrating representational geometry with layer-wise feature tracking. For the first time, it reveals systematic semantic feature shifts induced by SFT within sparse latent spaces and identifies layer-update patterns uniquely associated with safety alignment. The method precisely localizes key semantic features whose distributions are altered by SFT. All code and analyses are publicly released.
📝 Abstract
The cosine similarity between a large language model's hidden activations before and after Supervised Fine-Tuning (SFT) remains very high. This, at first glance, suggests that SFT leaves the model's activation geometry largely undisturbed. However, projecting both sets of activations through a Sparse Autoencoder (SAE) pretrained on the base model reveals that the underlying sparse latents diverge significantly. We introduce a novel investigative pipeline which utilizes these pretrained SAEs as a high-resolution diagnostic tool to mechanistically investigate the drivers of this representational divergence. Through our analytical pipeline, we discover task-specific and layer-specific distributions of the precise semantic features that are systematically altered during supervised fine-tuning. We additionally identify a layer-wise update profile specific to safety alignment. All code, experimental scripts, and analysis files associated with this work are publicly available at: https://github.com/ruhzi/sae-investigation.
Problem

Research questions and friction points this paper is trying to address.

Supervised Fine-Tuning
Sparse Autoencoder
Activation Geometry
Representational Divergence
Safety Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Supervised Fine-Tuning
Sparse Autoencoder
Mechanistic Interpretability
Activation Geometry
Representational Divergence
R
Ruhaan Chopra
Independent Researcher