Statistical Benefits of Fine-Tuning from Pretrained Initialization in Diagonal Linear Networks

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the theoretical opacity surrounding how fine-tuning effectively leverages pretrained weight information. By employing sparse linear regression and diagonal linear networks, this work conducts a dynamical analysis of gradient descent to elucidate how pretrained initialization influences implicit bias and sample complexity. It theoretically demonstrates that gradient-based fine-tuning implicitly exploits the pretrained support set, thereby achieving the statistical efficiency analogous to weighted Lasso. Furthermore, when the model correctly inherits the sign pattern of the pretrained parameters, fine-tuning significantly reduces the sample complexity required for target parameter recovery.
πŸ“ Abstract
Adapting pretrained models to downstream tasks with limited data has become a central paradigm in modern deep learning. Yet, despite its widespread practical success, how fine-tuning leverages information from pretraining remains poorly understood theoretically. We study fine-tuning from pretrained weights through the lens of sparse linear regression and two-layer diagonal linear networks. In our setting, pretraining provides information through the support (and signs) of the initialization predictor, which may contain coordinates relevant to the downstream task. We show how pretrained information reshapes the implicit bias and training dynamics, and can thereby reduce the sample complexity of recovering the target parameters and support. In particular, for a clean initialization with correctly inherited signs, we show that the required sample size is comparable to that of a weighted Lasso estimator that explicitly exploits the pretrained support through a suitably chosen regularizer. Our results thus show how information encoded in pretrained weights can be implicitly exploited by gradient-based fine-tuning, reducing the amount of data needed to recover a downstream task.
Problem

Research questions and friction points this paper is trying to address.

fine-tuning
pretrained models
implicit bias
sample complexity
sparse linear regression
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-tuning
pretrained initialization
diagonal linear networks
implicit bias
sample complexity
πŸ”Ž Similar Papers
No similar papers found.