🤖 AI Summary
This study addresses the long-standing challenge in continual learning where data attribution, catastrophic forgetting, and plasticity loss have been treated as isolated phenomena lacking a unified explanation. To bridge this gap, this work proposes a comprehensive theoretical framework that leverages token- and layer-wise update decomposition, Softmax force analysis, and residual connection decoupling to uncover the intrinsic relationship between local interactions and long-term learning dynamics. By explicitly distinguishing collision from erosion mechanisms, the framework unifies these phenomena as different manifestations of a single evolutionary process. Furthermore, it enables efficient data selection, interference control, and future learnability diagnostics. Ultimately, this research provides a generalizable perspective for continual learning that integrates both theoretical rigor and practical utility.
📝 Abstract
Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from what comes next. We show that these challenges are governed by the same evolving update--behavior interaction. We derive a token- and layer-wise decomposition of how learning from one token changes another prediction. By separating the softmax force, shared readout geometry, and residual connections, it exposes two interaction channels and yields a forward-computable approximation. Following this interaction through time reveals a unified picture of continual adaptation. Positive interaction identifies useful experience; negative interaction produces either concentrated collision or accumulated erosion; over longer horizons, updates reshape the shared geometry mediating future learning signals, reducing their transmission. These predictions lead to effective data selection, mechanism-specific controls for interference, and a readout-based diagnostic of future learnability whose degradation predicts the benefit of restoring the readout. Across models and training regimes, the same local interaction thus explains both what an update changes now and how learning today changes what can be learned tomorrow. This view connects data attribution, forgetting, and plasticity loss as distinct regimes of the same evolving learning dynamics.