🤖 AI Summary
This study addresses the unclear generalization mechanisms in auxiliary learning by developing an analytical nonlinear network fluctuation-dissipation theory within a teacher-student framework. We derive the online stochastic gradient descent (SGD) dynamical equations to systematically quantify the effects of task relatedness and gradient noise on generalization performance. By combining analytical solutions of differential equations with empirical validation, this work reveals the intrinsic relationship between main-auxiliary task errors and single-task errors, while elucidating the dynamical mechanism through which moderate gradient noise enhances generalization. Ultimately, this research provides a rigorous theoretical foundation for understanding implicit regularization effects in multi-task learning.
📝 Abstract
Auxiliary learning is an optimization paradigm in which a neural network's performance on a target task is improved by jointly training it on additional tasks. However, the mechanisms behind this improvement remain poorly understood. We study this problem using a teacher-student framework and derive a closed system of differential equations describing the dynamics of online stochastic gradient descent in the large-input limit. For linear networks, we obtain a closed-form expression for the generalization error to leading order in the learning rate, quantifying how task correlations and label noise determine the benefit of auxiliary learning. For non-linear activation functions, we develop a fluctuation-dissipation analytical theory that establishes a general relation linking the main and auxiliary errors to the corresponding single-task error. Numerical experiments support the theoretical predictions and show how auxiliary tasks improve generalization by balancing the forcing dynamics towards the optimal solution with gradient noise.