Normative Loss Landscape Navigation: A Trajectory-Based Approach to Mitigating Forgetting in Incremental Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses catastrophic forgetting in continual learning and the limitation of existing methods that overlook the Riemannian geometry of parameter space. By formulating the learning process as an optimal control problem over a curved loss landscape, this work proposes Trajectory-Modulated Loss Navigation (TMLN). Rather than relying on rigid Euclidean penalties, TMLN dynamically adjusts a diagonal empirical Fisher preconditioner using historical parameter trajectories. This mechanism directly safeguards critical directions within gradient updates, eliminating the need for auxiliary regularization terms. Experimental results demonstrate that TMLN significantly reduces inter-task loss barriers on both class-incremental and domain-incremental benchmarks, effectively mitigating catastrophic forgetting.
📝 Abstract
Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mitigate forgetting, current approaches are myopic. Conversely, standard regularization methods apply rigid, scalar Euclidean penalties that entirely ignore the underlying Riemannian geometry of the parameter space. To overcome this gap, we propose TMLN (Trajectory-Modulatory Landscape Navigation), a normative navigation policy that formalizes continual learning as an optimal control problem over a curved loss landscape. TMLN utilizes a memory-efficient diagonal empirical Fisher Information Matrix (FIM) to define a localized Riemannian manifold. To compensate for the spatial limitations of the diagonal approximation, TMLN dynamically modulates a preconditioner using the normalized historical trajectory of the network's parameter values. By integrating this trajectory-based preconditioning directly into the gradient update, we actively shield historically critical parameter directions without relying on additive penalties. Empirical evaluations on class- and domain-incremental benchmarks demonstrate that our method significantly reduces the loss barrier between consecutive tasks.
Problem

Research questions and friction points this paper is trying to address.

catastrophic forgetting
continual learning
incremental learning
loss landscape
Riemannian geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Learning
Riemannian Geometry
Fisher Information Matrix
Trajectory-Based Preconditioning
Optimal Control
🔎 Similar Papers
I
Isabelle Aguilar
School of Biomedical Engineering, University of Sydney
Z
Zayn Andre Zainal
School of Biomedical Engineering, University of Sydney
L
Luis Fernando Herbozo Contreras
School of Biomedical Engineering, University of Sydney
Z
Zhaojing Huang
School of Biomedical Engineering, University of Sydney
Omid Kavehei
Omid Kavehei
The University of Sydney
nanoelectronicsmedical electronicsaffective computinglearning machinesintegrated circuit design