🤖 AI Summary
This work addresses the fundamental problem of how task structure, learning rules, and input distribution jointly govern learning efficiency in nonlinear perceptrons—moving beyond conventional student–teacher frameworks and linear-output assumptions. We develop a stochastic-process-based flow equation that unifies the dynamical evolution of both supervised learning (SL) and reinforcement learning (RL). For the first time in nonlinear perceptrons, we quantitatively characterize the differential regulatory role of input noise: SL exhibits greater robustness to input noise, whereas RL suffers accelerated forgetting and slower task coverage. These theoretical predictions are empirically validated on MNIST, precisely reproducing learning/forgetting curves and cross-task interference magnitudes. Our framework establishes a novel, testable paradigm for analyzing nonlinear learning dynamics in both biological and artificial neural networks, providing a rigorous quantitative foundation for comparative analysis across learning paradigms.
📝 Abstract
The ability of a brain or a neural network to efficiently learn depends crucially on both the task structure and the learning rule. Previous works have analyzed the dynamical equations describing learning in the relatively simplified context of the perceptron under assumptions of a student-teacher framework or a linearized output. While these assumptions have facilitated theoretical understanding, they have precluded a detailed understanding of the roles of the nonlinearity and input-data distribution in determining the learning dynamics, limiting the applicability of the theories to real biological or artificial neural networks. Here, we use a stochastic-process approach to derive flow equations describing learning, applying this framework to the case of a nonlinear perceptron performing binary classification. We characterize the effects of the learning rule (supervised or reinforcement learning, SL/RL) and input-data distribution on the perceptron’s learning curve and the forgetting curve as subsequent tasks are learned. In particular, we find that the input-data noise differently affects the learning speed under SL vs. RL, as well as determines how quickly learning of a task is overwritten by subsequent learning. Additionally, we verify our approach with real data using the MNIST dataset. This approach points a way toward analyzing learning dynamics for more-complex circuit architectures.