fast-slow memory training

Design and evaluate training algorithms and memory mechanisms that split learning into fast and slow components (fast-adapting short-term memory plus slow-stable memory) to stabilize long-horizon rollouts and reduce temporal error accumulation; this includes devising architectures, update rules, loss functions, and metrics that preserve long-term dynamics during autoregressive or rollout-style sequence prediction.

fast-slowmemorytraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing work lacks a theoretical understanding of the high-performance mechanisms of state space models (SSMs), particularly regarding gradient optimization dynamics and long-range dependency learning. Method: We establish a gradient dynamical analysis framework, revealing for the first time the decisive role of memory capacity in guiding parameter update directions; prove the theoretical equivalence between S4 and its diagonalized variant; identify an intrinsic trade-off between memory length and precision; and propose an “initialization-as-design” paradigm—constructing fixed recurrent weight structures via controllable initialization, bypassing conventional adaptive updates. Contribution/Results: Theoretically and empirically, our approach significantly improves long-range memory modeling efficiency, achieving faster convergence and superior or competitive performance on language modeling and sequence forecasting tasks. It provides a novel, interpretable training pathway for SSMs.

Analyzing memory-accuracy tradeoff and theoretical equivalence in S4 modelsExplaining learning dynamics and memory mechanisms in state space modelsProposing improved initialization and fixed-weight training strategies

This work addresses the challenges of catastrophic forgetting and limited plasticity in large language models during parameter updates, as well as the insufficient performance of context-only learning. The authors propose a fast-slow learning framework that treats model parameters as “slow weights” to preserve general capabilities while optimizing the input context as “fast weights” to efficiently assimilate task-specific information. By integrating coordinated fast-slow weight training, context optimization, reinforcement learning, and KL divergence regularization, the method substantially mitigates forgetting in continual learning settings. Empirical results demonstrate up to a threefold improvement in sample efficiency, higher performance ceilings, and a 70% reduction in KL divergence compared to baseline approaches.

catastrophic forgettingcontinual learningin-context learning

Dynamic Dual Buffer with Divide-and-Conquer Strategy for Online Continual Learning

May 23, 2025
CD
Congren Dai
🏛️ Imperial College London | University of Electronic Science and Technology of China

To address severe catastrophic forgetting and inefficient memory updating in online continual learning (OCL), this paper proposes a Dynamic Dual-Cache Memory Framework. It comprises a short-term cache capturing instantaneous changes in streaming data and a long-term cache partitioned into multiple sub-buffers, where knowledge is archived via class prototypes. We innovatively leverage optimal transport theory to guide prototype-aware sample retention and integrate K-means clustering to preserve semantic richness. Furthermore, we design a Divide-and-Conquer memory updating strategy (DAC), decomposing global optimization into parallelizable subproblems. Evaluated on standard and class-imbalanced OCL benchmarks, our method achieves state-of-the-art performance: it significantly reduces forgetting rates and cuts memory update computational overhead by 42%.

Addresses catastrophic forgetting in online continual learningIntroduces dual memory system for dynamic and enduring knowledgeProposes Divide-and-Conquer strategy to optimize memory updates

This work addresses catastrophic forgetting in continual learning under non-stationary data streams by proposing the COLD framework, which introduces, for the first time, the Drift-Plus-Penalty stochastic optimization method from control theory into this domain. COLD formulates forgetting as a controlled dynamic process, employing virtual queues to track performance deviations on historical tasks and jointly minimizing the current task loss and queue drift at each optimization step. This mechanism explicitly governs the stability-plasticity trade-off. The framework provides theoretical guarantees on stability and convergence, and achieves significantly superior performance over state-of-the-art methods on standard benchmarks, enabling controllable and efficient suppression of catastrophic forgetting.

catastrophic forgettingcontinual learningnonstationary data streams

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional continual learning, which overly emphasizes preserving old knowledge to approximate joint training while neglecting real-time adaptation to new tasks in non-stationary environments. The problem is reformulated as an online optimization framework, with average lifelong error as the performance metric, and a notion of transfer efficiency is introduced to characterize the trade-off between stability and transient error induced by historical knowledge. Drawing on critical task duration theory, the study identifies conditions under which past knowledge shifts from beneficial to detrimental. Building on this insight, the paper proposes a predictive continual learning paradigm that dynamically models future tasks to optimize expected performance. Integrating online learning, transfer efficiency analysis, and convergence theory, an algorithm based on task-window interpolation is developed and validated on image classification and reinforcement learning benchmarks, demonstrating significant superiority over both joint training and independent learning, especially under distribution shift.

AdaptationCatastrophic ForgettingContinual Learning

This work addresses the inefficiency of autoregressive rollout generation in reinforcement learning post-training and the inability of existing speculative decoding methods to adapt to continuously evolving policies. To overcome these limitations, we propose SpecRoll—a dual-timescale adaptive speculative rollout engine that significantly accelerates generation while strictly preserving the target policy’s sampling distribution. SpecRoll employs a lightweight future-token head for parallel proposal generation and integrates a backpropagation-free Reflex module for local hidden-state correction along trajectories. It further incorporates concurrency-aware sparse-tree verification, exact target validation, adaptive fast-slow path routing, and delayed validator feedback. Evaluated across models ranging from 1.5B to 14B parameters and three mathematical reasoning benchmarks, SpecRoll achieves 1.26–2.15× faster rollout generation and 1.21–2.04× end-to-end speedup, consistently outperforming FastGRPO.

efficiency bottleneckpolicy adaptationreinforcement learning

This work investigates the convergence behavior of stochastic gradient descent (SGD) operating near the edge of stability under large learning rates, addressing a gap in existing theory that lacks rigorous analysis in stochastic settings. Focusing on linear classifiers and two-layer neural networks trained with multiclass cross-entropy loss, the study integrates tools from stochastic optimization theory, dynamical systems analysis, and expected loss control to uncover an alternating mechanism between curvature-driven oscillations and stable descent. The paper establishes the first rigorous convergence guarantees for large-learning-rate SGD, demonstrating that its intrinsic stochasticity induces self-stabilization and enables optimal convergence rates within a fixed number of iterations. These theoretical findings are corroborated by experiments, highlighting the efficacy and superiority of large-step SGD.

convergenceedge of stabilitylarge learning rates

This work addresses the catastrophic performance degradation in 4-bit quantized large language models (LLMs) operating in dynamic financial environments, where continual memory editing induces severe knowledge forgetting. To mitigate this issue, the authors propose a stability-aware memory editing framework that, for the first time, integrates stability mechanisms into the memory update process of quantized LLMs. The approach combines low-rank adaptation (LoRA), domain-specific priority modulation tailored to finance, and closed-loop stability control, augmented with a novel degradation-debt tracking mechanism and adaptive editing intensity strategy. Evaluated on a corpus of 88,021 UK financial documents, the method reduces knowledge degradation by 11–17% compared to baseline approaches and improves test-time generalization success rates by six percentage points, achieving up to 28%.

catastrophic forgettingfinancial domainmemory editing

Hot Scholars

KS

Kaustubh Sharma

IIT Roorkee
Mechanistic InterpretabilityDiffusion ModelsAI Safety
AG

Anshul Gupta

Research Assistant at Idiap Research Institute; PhD candidate at EPFL
computer visionmulti-modal learningmachine learning
MM

Michele Magno

ETH Zurich
Wireless sensor networksSmart Sensors and Internet of ThingsWake up RadioPower management
XC

Xilong Cheng

Communication University of China
Role-Playing AgentsComputer visionTime Series Forecasting
SC

Seokeon Choi

Qualcomm AI research
Computer visionMachine learningImage generationDomain generalization