🤖 AI Summary
This study addresses catastrophic forgetting and the difficulty of reusing knowledge for novel task combinations in continual learning by proposing the GaTaP algorithm. This method introduces a bilevel optimization framework that achieves knowledge recomposition through a gating variable selection module. It learns gating parameters via closed-form solutions, synchronizing local error signals for both gating adaptation and weight updates, while optimizing convolutional neural networks through target propagation. Experimental results demonstrate that GaTaP significantly improves the retention of previously learned tasks in class-incremental learning and enables few-shot combinatorial generalization to unseen tasks. Furthermore, the findings reveal the capacity of the gating mechanism to capture reusable task structures.
📝 Abstract
Continual learning is typically framed as acquiring new knowledge without catastrophically forgetting previous tasks. However, a flexible continual learner should also be able to reuse and recombine previously acquired knowledge to rapidly solve novel task compositions. We introduce Gated Target Propagation (GaTaP), a continual learning algorithm in which task-specific gating variables---learned through a closed-form inner loop update---selectively suppress or enhance network modules. Network parameters are learned in a slower timescale outer loop, using the same local difference target propagation error signal as is used for adapting gating variables. We provide tractable experiments on class-incremental learning scenarios for both multilayer perceptron and convolutional network architectures. We show strong performance retention on previously learned tasks, as well as compositional generalization to unseen tasks, achieved through few-shot gain adaptation at inference. We analyze learned gating patterns and find that related tasks exhibit similar gating patterns, suggesting that inferred gates capture meaningful, reusable task structure. Overall, GaTaP provides a powerful framework for jointly ameliorating catastrophic forgetting and enabling few-shot compositional generalization in neural network models.