π€ AI Summary
This work investigates the nature of catastrophic forgetting in continual learning, disentangling the effects of representation loss from interface drift. By splicing the early layers of a pre-trained model with the later layers of a sequentially updated model and introducing task-specific βtransfer keysβ to align internal interfaces, the method effectively recovers forgotten knowledge. Combining anchor activation pairing with a compact interface alignment operator, the approach demonstrates on ResNet and small Vision Transformers that forgetting primarily stems from interface drift rather than irreversible loss of learned representations, as latent features remain accessible. Experiments on benchmarks such as Split CIFAR-100 show that most of the original performance can be restored, highlighting the critical role of re-indexing latent computations for knowledge recovery.
π Abstract
Catastrophic forgetting is often framed as a representational problem: after sequential training, a model appears to lose the features that supported performance on earlier tasks. We challenge the stronger form of this view. Across controlled continual-learning settings, we find that a significant portion of apparent forgetting can be attributed to interface drift between internal stages rather than permanent erasure of task-relevant computation. We study this phenomenon through a stitched evaluation protocol that combines early computation from a post-update network with late computation from its predecessor, optionally mediated by a compact, task-specific transport key. We describe transport keys at a systems level as compact interface-alignment operators estimated from a small set of paired anchor activations and evaluated through model stitching. On split CIFAR-100 with a ResNet-style network, transport keys recover most of the original Task A performance after sequential training on Task B. On a compact vision transformer, we observe a similar recovery pattern. These results suggest that continual learning may require better mechanisms for indexing and re-accessing latent computations, not only methods that prevent weight change.