🤖 AI Summary
This study aims to elucidate the intrinsic mechanisms by which pre-training accelerates downstream transfer learning. Leveraging a random mapping memorization task, we identify and analyze two distinct transfer regimes: equivalent and non-equivalent transfer. For the first time, we decompose transfer effects into final-layer magnitude-driven and structure-driven components, thereby clarifying previously counterintuitive transfer phenomena. Furthermore, through systematic ablation studies, we quantify the specific contributions of inter-layer covariance and weight changes across network layers to overall transfer performance. By comprehensively revealing the underlying principles governing transfer dynamics, this work provides critical theoretical guidance for designing principled and effective pre-training strategies.
📝 Abstract
A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a "trivial" magnitude-driven transfer in the last layer, and a "non-trivial" structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.