When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
This study investigates when the training history of neural networks better predicts future learning than their current state, a dimension frequently overlooked in research on plasticity loss and critical periods. Employing short-horizon response probing, synthetic regression-based predictor screening, and statistical consistency measures such as the intraclass correlation coefficient (ICC), this work systematically compares the predictive capacity of training history versus current states for small multilayer perceptrons. It provides the first quantification of the predictive gain offered by training history relative to current states, delineating its temporal boundaries. The findings reveal that historical information confers a significant predictive advantage only when the current state lacks informativeness, such as during early training stages, with no discernible difference observed later. Consequently, this work establishes the novel insight that training history becomes valuable precisely when the current state is uninformative.