🤖 AI Summary
In online continual learning, conventional experience replay often neglects inter-sample spatial correlations, leading to biased risk estimation and catastrophic forgetting. This work proposes SPHERE, a novel method that introduces the first representation-kernel-based risk aggregation mechanism to suppress isolated noise while preserving consistently high-risk samples. Furthermore, SPHERE formulates replay allocation as an entropy-regularized optimal transport problem, thereby achieving a robust memory sampling strategy. Extensive evaluations demonstrate that the proposed approach significantly improves accuracy and effectively mitigates forgetting rates across diverse tasks, including visual recognition, instruction tuning, and code generation.
📝 Abstract
Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated responses. We introduce SPatial coHErent risk control for REplay (SPHERE), a general replay-allocation method applicable across a broad range of learning settings. SPHERE uses a representation kernel to aggregate signed prospective loss changes, attenuating unsupported spikes while retaining coherent increases. It then formulates allocation as entropy-regularized transport, redistributing uniform source mass toward supported high-risk regions while penalizing long-distance transfers. We derive replay coefficients from the transport objective's sensitivity to the original loss changes and blend them with uniform replay to maintain baseline rehearsal. Our analysis establishes conditions under which kernel aggregation improves risk estimation and bounds transport-value inflation due to residual noise and smoothing bias. Experiments demonstrate that SPHERE improves accuracy and reduces forgetting across noisy-label vision tasks, continual language-model instruction tuning, and code-generation reinforcement learning with incomplete test rewards.