🤖 AI Summary
This study investigates the impact of task granularity ordering on catastrophic forgetting in continual learning, presenting the first systematic evaluation of three learning strategies—coarse-to-fine, fine-to-coarse, and flat learning—on CIFAR-100. Leveraging Elastic Weight Consolidation (EWC), the authors assess model performance using accuracy, F1 score, and continual learning–specific metrics to quantify the retention of previously acquired knowledge. The findings demonstrate that initializing learning with coarse-grained categories before introducing fine-grained tasks significantly mitigates catastrophic forgetting and enhances backward transfer. This suggests that incorporating hierarchical priors to construct stable representations offers an effective principle for designing task sequences in incremental learning scenarios, thereby providing a novel strategy for optimizing continual learning systems.
📝 Abstract
Catastrophic forgetting, the abrupt loss of previously acquired knowledge upon learning new information, remains the central challenge in Continual Learning. This project investigates whether the order in which a model learns information affects how well it retains knowledge. Specifically, we ask: does learning general categories first (like "animals" vs "vehicles") before learning specific classes (like "dog" vs "cat") reduce forgetting compared to learning all classes at once?
We test three approaches on CIFAR-100: (1) Coarse-to-Fine: train on 2 super-classes, then expand to 10 specific sub-classes, (2) Fine-to-Coarse: train on 10 sub-classes, then group into 2 super-classes, and (3) Flat: train on all 10 classes from the start. We use Elastic Weight Consolidation (EWC) to prevent forgetting during transitions. Our hypothesis is that learning general patterns first creates a stable foundation that helps the model retain knowledge when learning more detailed distinctions. We evaluate using standard metrics (accuracy, precision, recall, F1) plus continual learning metrics like backward transfer and forgetting rates. This work could inform how we design learning sequences for real-world systems that need to learn incrementally.