Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析注意力模式和层激活来探索微调如何改变大型语言模型的内部表示,并检查这些变化与任务相关组件的关系。
📝 Abstract
Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.
Problem

Research questions and friction points this paper is trying to address.

fine-tuning
internal representations
task-relevant components
EAP-identified components
cross-task performance transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-tuning
internal representations
EAP-identified components
functional localisation
cross-task performance
🔎 Similar Papers
No similar papers found.