🤖 AI Summary
Although LoRA fine-tuning is widely adopted for large language models, the mechanisms by which it alters internal representational structures remain poorly understood. This work proposes a delta activation framework, integrated with sparse autoencoders (SAEs), to systematically isolate and analyze adapter-specific representations introduced by LoRA within the residual stream. Leveraging cosine similarity, principal angles, and centered kernel alignment (CKA), the study reveals that LoRA-induced feature dictionaries exhibit weak geometric alignment with pre-trained features, while adapter-specific SAEs achieve superior reconstruction performance. Feature density increases with both rank and network depth, yet geometric discrepancies remain stable across different ranks, suggesting that fine-tuning may generate novel representational structures not readily captured by existing interpretability tools.
📝 Abstract
Low-Rank Adaptation (LoRA) has emerged as a widely adopted approach for adapting large language models, yet the internal representational changes induced by LoRA fine-tuning remain insufficiently understood. In this work, we investigate the geometry of LoRA-induced representations using Sparse Autoencoders (SAEs). We introduce a delta activation framework that isolates the adapter-specific contribution to the residual stream.
Using Gemma-2-9B with LoRA ranks 4, 8, 16, and 32, we train adapter-specific SAEs across multiple transformer layers and compare their learned feature spaces with pretrained SAE dictionaries. We evaluate representational alignment using cosine similarity between decoder directions, principal-angle analysis of feature subspaces, and Centered Kernel Alignment (CKA) between activation representations.
Across layers and ranks, we consistently observe comparatively weak geometric alignment between LoRA-induced feature dictionaries and pretrained SAE features. Adapter-specific SAEs also reconstruct delta activations more effectively than pretrained SAEs, suggesting that LoRA updates occupy partially distinct representational structure within the residual stream. Additionally, feature density increases with rank and depth, while geometric divergence remains relatively stable across ranks.
These findings provide empirical evidence that LoRA fine-tuning can induce feature structures that are not fully captured by pretrained interpretability dictionaries, with implications for mechanistic interpretability, adaptation analysis, and safety auditing of fine-tuned language models.