Layer-Wise Evolution of Representations in Fine-Tuned Transformers: Insights from Sparse AutoEncoders

📅 2025-02-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the hierarchical representation evolution of pretrained Transformers (e.g., BERT) during fine-tuning—specifically, how models balance preservation of general-purpose features against acquisition of task-specific ones. We propose an activation interpretability framework based on sparse autoencoders (SAEs), integrating inter-layer similarity metrics, token-level activation visualization, and cross-dataset comparative experiments. Our analysis systematically reveals, for the first time, a universal layer-wise functional specialization pattern: early layers retain generic linguistic features, middle layers exhibit progressive transitional behavior, and deeper layers specialize in task-adaptive representations. This finding provides both theoretical grounding and empirical evidence for controllable fine-tuning, model compression, and interpretable AI.

Technology Category

Machine Learning: Representation LearningNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsComputer Vision: Representation Learning for Vision

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Web query analysis, representation and understandingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Fine-tuning pre-trained transformers is a powerful technique for enhancing the performance of base models on specific tasks. From early applications in models like BERT to fine-tuning Large Language Models (LLMs), this approach has been instrumental in adapting general-purpose architectures for specialized downstream tasks. Understanding the fine-tuning process is crucial for uncovering how transformers adapt to specific objectives, retain general representations, and acquire task-specific features. This paper explores the underlying mechanisms of fine-tuning, specifically in the BERT transformer, by analyzing activation similarity, training Sparse AutoEncoders (SAEs), and visualizing token-level activations across different layers. Based on experiments conducted across multiple datasets and BERT layers, we observe a steady progression in how features adapt to the task at hand: early layers primarily retain general representations, middle layers act as a transition between general and task-specific features, and later layers fully specialize in task adaptation. These findings provide key insights into the inner workings of fine-tuning and its impact on representation learning within transformer architectures.
Problem

Research questions and friction points this paper is trying to address.

Analyzing fine-tuning mechanisms in transformers
Exploring feature adaptation across BERT layers
Understanding task-specific vs. general representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuning pre-trained transformers
Analyzing activation similarity
Visualizing token-level activations
🔎 Similar Papers
No similar papers found.