🤖 AI Summary
This work addresses the theoretical gap in partial parameter reuse within transfer learning. We investigate the conditions under which transferring only a subset of parameters from a pretrained ReLU convolutional neural network remains effective—and when it fails—for downstream tasks. By establishing a rigorous theoretical link between upstream feature learning capability and downstream performance, we derive a discriminative criterion for parameter transferability, identifying key determinants: feature generality, task alignment, and subspace structure of the transferred parameters. Our analysis reveals that transfer degrades performance below from-scratch training when inherited parameters cannot support discriminative feature representations required by the downstream task. Combining formal derivation with numerical experiments, we validate both the plausibility of our knowledge-transfer-path modeling and the empirical validity of the proposed conditions. This yields the first systematic, theoretically grounded framework for controllable and interpretable partial-parameter transfer.
📝 Abstract
Parameter transfer is a central paradigm in transfer learning, enabling knowledge reuse across tasks and domains by sharing model parameters between upstream and downstream models. However, when only a subset of parameters from the upstream model is transferred to the downstream model, there remains a lack of theoretical understanding of the conditions under which such partial parameter reuse is beneficial and of the factors that govern its effectiveness. To address this gap, we analyze a setting in which both the upstream and downstream models are ReLU convolutional neural networks (CNNs). Within this theoretical framework, we characterize how the inherited parameters act as carriers of universal knowledge and identify key factors that amplify their beneficial impact on the target task. Furthermore, our analysis provides insight into why, in certain cases, transferring parameters can lead to lower test accuracy on the target task than training a new model from scratch. Numerical experiments and real-world data experiments are conducted to empirically validate our theoretical findings.