๐ค AI Summary
This study addresses the performance degradation in continual learning caused by recency bias during sequential task updates. Working within an overparameterized linear regression framework, it quantifies the excess loss incurred by sequential exact fitting and demonstrates that distribution-level forgetting and population loss converge to the same limit. The authors introduce the theoretical concept of โsequential priceโ and derive its analytical solution under Elastic Weight Consolidation (EWC) regularization. Through experiments on Jester and Rotated MNIST benchmarks, this work validates the predicted values of the sequential price and confirms its decay as EWC regularization strength increases. The strong alignment between theoretical predictions and empirical results establishes an interpretable theoretical foundation for mitigating catastrophic forgetting in continual learning scenarios.
๐ Abstract
Sequential task updates are fundamental to continual learning, but their recency bias can impose a lasting performance cost. We study this cost in an overparameterized linear-regression model with i.i.d. task sampling. We prove that distribution-level forgetting and population loss converge to the same stationary limit. This common limit separates exactly into the intrinsic loss asymptotically attained by joint training and an additional sequential price, and in more homogeneous task geometries the two terms coincide, making the total loss twice that of joint training. We further analyze fixed-strength elastic weight consolidation (EWC) under general task curvatures and characterize its stationary sequential price at every regularization strength. Under strong regularization, the price decays inversely with EWC strength while convergence to stationarity slows at the same scale. On the Jester joke-rating dataset, the theory exactly quantifies both the sequential price generated by naturally conflicting user preferences and its reduction by EWC.