🤖 AI Summary
This study addresses the unclear theoretical mechanisms underlying few-shot fine-tuning, particularly how pre-training facilitates new feature learning. We investigate a two-timescale fine-tuning strategy for two-layer ReLU networks under Gaussian multi-index models to elucidate the advantages of implicit bias induced by pre-training. Our theoretical analysis demonstrates that the proposed algorithm preserves previously acquired features while efficiently learning new ones, thereby overcoming the performance bottlenecks of random initialization in low-data regimes. Specifically, we prove that the target parameters can be recovered using only O(d) samples, significantly outperforming randomly initialized baselines. This work provides rigorous theoretical foundations for understanding the pre-training and fine-tuning paradigm in deep learning.
📝 Abstract
Fine-tuning pre-trained models on specialized tasks with scarce data is central to modern deep learning. Despite its empirical success, theoretical understanding of fine-tuning remains limited. We introduce a Gaussian multi-index setting to study fine-tuning from pre-trained weights, where the teacher network has $m+1$ features, $m$ of which are learned during pre-training and one of which must be learned during fine-tuning. For two-layer ReLU networks, we show that two-timescale training, i.e., updating the outer weights infinitely faster than the hidden ones, learns the new task-specific feature while preserving the pre-trained ones in the model representation. Moreover, only $\mathcal{O}(d)$ fine-tuning samples are required for this recovery, independently of the number of pre-trained features. In contrast, with random initialization, the same number of samples is insufficient to recover the target parameters. Our results therefore demonstrate that pre-training can induce an implicit bias with a clear statistical advantage over random initialization, enabling feature learning from scarce fine-tuning data.