Overcoming catastrophic forgetting in neural networks

📅 2025-07-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses catastrophic forgetting in continual learning by proposing and systematically evaluating an improved implementation of Elastic Weight Consolidation (EWC). We reproduce EWC on the PermutedMNIST and RotatedMNIST benchmarks, augmenting the analysis with dropout integration and comprehensive hyperparameter sensitivity studies. Results demonstrate that EWC effectively balances knowledge retention and new-task adaptation: it substantially mitigates forgetting—outperforming unregularized SGD baselines—and enhances overall continual learning stability and generalization, albeit with a modest reduction in convergence speed on new tasks. Our study confirms EWC’s robustness across canonical non-stationary task sequences and reveals synergistic effects between regularization strength and architectural choices (e.g., dropout rate) on continual learning efficacy. Crucially, this work provides reproducible empirical evidence supporting lightweight, regularization-based continual learning approaches.

Technology Category

Machine Learning: Life-Long and Continual LearningSearch and Optimization: Learning to SearchIntelligent Robots: Learning & Optimization for ROB

Application Category

Economics, Online Markets and Human Computation: Sustainability of Web economicsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingWeb Mining and Content Analysis: Robustness and generalizability of Web mining methods
📝 Abstract
Catastrophic forgetting is the primary challenge that hinders continual learning, which refers to a neural network ability to sequentially learn multiple tasks while retaining previously acquired knowledge. Elastic Weight Consolidation, a regularization-based approach inspired by synaptic consolidation in biological neural systems, has been used to overcome this problem. In this study prior research is replicated and extended by evaluating EWC in supervised learning settings using the PermutedMNIST and RotatedMNIST benchmarks. Through systematic comparisons with L2 regularization and stochastic gradient descent (SGD) without regularization, we analyze how different approaches balance knowledge retention and adaptability. Our results confirm what was shown in previous research, showing that EWC significantly reduces forgetting compared to naive training while slightly compromising learning efficiency on new tasks. Moreover, we investigate the impact of dropout regularization and varying hyperparameters, offering insights into the generalization of EWC across diverse learning scenarios. These results underscore EWC's potential as a viable solution for lifelong learning in neural networks.
Problem

Research questions and friction points this paper is trying to address.

Addresses catastrophic forgetting in neural networks during continual learning
Evaluates Elastic Weight Consolidation performance in supervised learning benchmarks
Analyzes trade-offs between knowledge retention and new task adaptability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elastic Weight Consolidation reduces catastrophic forgetting
Evaluated EWC with PermutedMNIST and RotatedMNIST benchmarks
Analyzed EWC's balance between retention and adaptability
🔎 Similar Papers
No similar papers found.
B
Brandon Shuen Yi Loke
École Polytechnique Fédérale de Lausanne
F
Filippo Quadri
École Polytechnique Fédérale de Lausanne
G
Gabriel Vivanco
École Polytechnique Fédérale de Lausanne
M
Maximilian Casagrande
École Polytechnique Fédérale de Lausanne
S
Saúl Fenollosa
École Polytechnique Fédérale de Lausanne