Fast Learning Rate Transfer in Shallow Linear Networks at Growing Training Horizons

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear impact of training duration on learning rate transfer mechanisms as model width increases. By analyzing shallow linear networks under full-batch gradient descent with training steps $T=o(\sqrt{n})$, this work investigates the conditions for rapid hyperparameter transfer. Leveraging spectral assumption analysis and random matrix theory, it provides the first characterization of how finite-width perturbation scales and local curvature influence transferability, while deriving the limiting distribution of optimal hyperparameters. The results demonstrate that short-duration training enables fast learning rate transfer and reveal the dominant role of extreme eigenvalues of the data Gram matrix in this process. These findings establish a theoretical foundation for hyperparameter transfer in wide networks.
📝 Abstract
Hyperparameter transfer across model width can substantially reduce the cost of tuning large neural networks, but its behavior when the training horizon grows with width is not fully understood. Building on the framework of fast hyperparameter transfer (Ghosh et al., 2026), which formalizes when transfer is effective, we investigate conditions that ensure fast transfer in the growing-horizon regime. Specifically, we study learning-rate transfer in a shallow linear network with a single trainable hidden matrix, trained by full-batch gradient descent. Under additional spectral assumptions, our main results are threefold. (i) We prove fast learning-rate transfer as $n,T\to\infty$ whenever $T=o(\sqrt{n})$. (ii) We characterize the transfer rates through the finite-width perturbation scale, the first-order sensitivities of the loss and its learning-rate derivative to finite-width perturbations, and the local loss curvature. (iii) We derive limiting distributions for the optimal learning rate and optimized loss, governed by fluctuations associated with the extreme eigenvalues of the data Gram matrix. These results clarify how spectral structure and local loss sensitivities govern learning-rate transfer at growing horizons.
Problem

Research questions and friction points this paper is trying to address.

hyperparameter transfer
learning rate transfer
shallow linear networks
growing training horizon
Innovation

Methods, ideas, or system contributions that make the work stand out.

Learning Rate Transfer
Shallow Linear Networks
Growing Training Horizon
Spectral Structure
Hyperparameter Transfer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mana Sakai
The University of Tokyo; RIKEN Center for Advanced Intelligence Project
Masaaki Imaizumi
Masaaki Imaizumi
The University of Tokyo / RIKEN AIP
StatisticsMachine Learning