One-Step Curvature Probes Miss the Fitting Operator: Retained Capacity and Terminal Null-Space Correction for Continual Learning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of single-step curvature probing in accurately evaluating terminal forgetting within continual learning. Leveraging an overparameterized linearized framework, we analyze the displacement inflation effect of projected gradient descent, revealing that the terminal quadratic cost is jointly determined by preserved fitting capacity, endpoint directional curvature, and null-space components. Accordingly, we propose a terminal null-space correction method that decouples the curvature-optimal fitting lower bound from algorithm-dependent null-space excess, effectively rectifying the bias inherent in conventional single-step probing. Experiments on Permuted MNIST and Split CIFAR-100 demonstrate that this correction significantly reduces old-task loss, validating the multi-factor dependency mechanism underlying the terminal cost.
📝 Abstract
A one-step curvature probe evaluates an initial direction, whereas continual learners are judged after reaching comparable new-task fit. In an overparameterized linearization, projected gradient descent converges to $\Delta_P=PJ^\top(JPJ^\top)^{-1}r$, and its squared-displacement inflation is exactly the reciprocal of the retained fitting capacity $c_P(r)$. More generally, the terminal old-task quadratic ratio factorizes as $1/[c_P(r)G_{\rm end}(P,r)]$, where $G_{\rm end}$ compares curvature along endpoint directions. In the rank-one case, $G_{\rm end}$ equals the one-step probe gain; for multiple outputs, the two gains can differ. The terminal quadratic also separates into a curvature-optimal fitting floor and an algorithm-dependent null-space excess, motivating terminal null-space correction, which preserves linearized new-task outputs on its Jacobian batch. Controlled checks validate the local quadratic and show that projection can greatly reduce matched-norm curvature while barely changing terminal forgetting. Across common-threshold configurations, projection yields the larger signed old-task loss change in 73/99 matched pairs, retained fitting capacity falls with rank, and fixed-rank comparisons separate terminal forgetting even when probe gain is approximately matched. On Permuted MNIST and Split CIFAR-100, terminal correction decreases signed old-task loss in 166/180 method--dataset--seed pairs under the all-seed intention-to-correct analysis; because acceptance and outcome reporting use the same held-out split, this result is descriptive and test-conditioned. Overall, terminal cost depends jointly on retained fitting capacity, endpoint-direction curvature, and the null-space component selected by optimization.
Problem

Research questions and friction points this paper is trying to address.

continual learning
catastrophic forgetting
curvature probe
retained fitting capacity
terminal null-space
Innovation

Methods, ideas, or system contributions that make the work stand out.

continual learning
retained fitting capacity
terminal null-space correction
curvature probe
catastrophic forgetting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Abu Sa-Adat Mohamed Moon-Im Al Ahsan
Department of Computer Science and Engineering, BRAC University, Bangladesh
Ibne Farabi Shihab
Ibne Farabi Shihab
Iowa State University
Deep LearningroboticsLarge Language Model
M
Md Najmus Swaqeeb
BRAC University, Bangladesh