input-driven forgetting

Designs and implements mechanisms that adaptively reduce or refresh the influence of prior latent states in response to incoming inputs, including exponential-decay schedules, input-driven adaptive forgetting rules, learnable latent-forgetting gates, and stale-state refresh gating. Analyzes and tunes these mechanisms to preserve consistency and stability of internal representations during sustained maneuvers or periods of stale observations and to trigger state refreshes when input patterns indicate hazards or significant change.

input-drivenforgetting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a fundamental limitation in existing adaptive methods, which treat environmental non-stationarity—particularly drift—as mere noise or distributional shift, thereby overlooking the progressive loss of organizational coherence between system and environment over time. To overcome this, the paper introduces the principle of Egregious Drift Regulation (EDR), reframing drift as a regulatory signal of coherence mismatch. EDR enables long-term coherent adaptation by dynamically adjusting the system’s internal structure to maintain, reorganize, or transition its operational mechanisms. Departing from conventional error-minimization objectives, this approach shifts the adaptive goal toward coherence regulation, integrating adaptive control with embodied cognition theory. It realizes a mechanism-centered “emergent machine” architecture that unifies state mechanisms, attractor dynamics, coherence metrics, reconfiguration dynamics, and cross-mechanism memory. The resulting framework offers a principled solution for intelligent systems operating in persistently non-stationary environments, substantially enhancing their long-term functional coherence.

adaptive systemsdriftenactive cognition

This study addresses the lack of a unified dynamical explanation for selective retention mechanisms in neural networks by proposing a theoretical framework termed "repeated reinforcement and persistent forgetting." Departing from the conventional view that treats forgetting as a deficiency, this work establishes it as a controllable inductive bias. Grounded in an independent feature model and small-step approximation, a three-tiered theoretical derivation reveals that forgetting induces spectral filtering and minimum description length (MDL)-style compression effects, with conclusions holding irrespective of architecture or scale. Experiments validate the joint selective mechanism of reinforcement and forgetting, demonstrating that forgetting exerts a decisive influence on both the composition of the retained set and its temporal dynamics.

forgetting dynamicsinductive biasselective retention

Time-Scale Coupling Between States and Parameters in Recurrent Neural Networks

Aug 16, 2025
LL
Lorenzo Livi
🏛️ OPIT – Open Institute of Technology

Despite their widespread success, the mechanisms underlying the superior trainability and stability of gated RNNs remain poorly understood—particularly how fixed global learning rates yield effective optimization. Method: We theoretically analyze the implicit adaptive learning rate behavior induced by gating mechanisms, deriving exact Jacobian matrices for leaky-integrator neurons and gated RNNs. Using first-order expansions, we characterize how scalar and multidimensional gates modulate gradient propagation, effective step sizes, and anisotropic parameter updates by coupling temporal scales in state space with update dynamics in parameter space. Contribution/Results: We establish that gating units act not only as memory controllers but also as data-driven preconditioners, spontaneously exhibiting optimizer-like properties—including learning-rate scheduling, momentum, and Adam-style adaptation—without explicit algorithmic design. Experimental validation confirms that the resulting gradient corrections, though small, are persistent and effective. This work provides the first systematic theoretical explanation for the robust training dynamics of gated RNNs.

Coupling between state-space time scales and parameter dynamicsGates as data-driven preconditioners in optimization trajectoriesHow gating mechanisms in RNNs induce adaptive learning-rate behavior

This work addresses the issue of optimizer state stagnation in large language model pretraining under low-precision quantization, where rounding errors impair the adaptivity of quantized optimizer states. The study systematically analyzes the quantization behavior of low-precision exponential moving average (EMA) states and proposes a reset-based mechanism to restore optimizer responsiveness. A predictive model of state stagnation is developed to elucidate the efficacy of the reset strategy, leading to a theoretically grounded derivation of the optimal reset interval. Through a combination of quantization analysis, stochastic process modeling, and controlled experiments, the approach is validated both in simulation and real-world LLM pretraining, demonstrating substantial recovery of performance loss while significantly reducing optimizer memory overhead.

LLM pre-traininglow-precision trainingoptimizer states

Sufficient conditions for offline reactivation in recurrent neural networks

May 22, 2025
NH
Nanda H Krishna
🏛️ Mila - Quebec AI Institute | Université de Montréal | McGill University | CIFAR

Whether noise-driven recurrent neural networks (RNNs) can autonomously replay task-evoked neural activity during input-free resting periods remains an open question—particularly whether task-optimized networks inherently possess offline reactivation capability. Method: We formulate the network dynamics via stochastic differential equations, establish Lyapunov stability conditions, and validate our theory numerically on spatial localization and head-direction estimation tasks. Contribution/Results: We derive the first rigorous mathematical sufficient condition for offline reactivation in RNNs. We prove that denoising dynamics—enabling faithful replay—naturally emerge from smooth stimulus encoding and change-driven optimization, without ad hoc mechanisms. Both theoretical analysis and numerical experiments demonstrate that networks satisfying these optimization principles spontaneously recapitulate online activity patterns during rest, achieving reactivation fidelity exceeding 92%. This reveals offline reactivation as an intrinsic, emergent property of optimally trained recurrent systems, bridging online computation and offline memory consolidation.

Conditions for neural reactivation in task-optimized networksMathematical framework for reactivation in noisy recurrent circuitsValidation via spatial and head direction estimation tasks

Latest Papers

What's happening recently
View more

This work addresses catastrophic forgetting in continual learning under non-stationary data streams by proposing the COLD framework, which introduces, for the first time, the Drift-Plus-Penalty stochastic optimization method from control theory into this domain. COLD formulates forgetting as a controlled dynamic process, employing virtual queues to track performance deviations on historical tasks and jointly minimizing the current task loss and queue drift at each optimization step. This mechanism explicitly governs the stability-plasticity trade-off. The framework provides theoretical guarantees on stability and convergence, and achieves significantly superior performance over state-of-the-art methods on standard benchmarks, enabling controllable and efficient suppression of catastrophic forgetting.

catastrophic forgettingcontinual learningnonstationary data streams

This work addresses the challenge of non-stationary reinforcement learning scenarios where task identifiers or environmental change signals are absent, a setting in which existing deep reinforcement learning methods often struggle to adapt effectively. To overcome this limitation, the paper introduces Space-sampled Value Decay (SVD), an explicit forgetting mechanism inspired by biological memory processes. SVD can be seamlessly integrated into value-based algorithms such as DQN and SAC, enabling dynamic adjustment of value functions without requiring prior knowledge of environmental shifts. Empirical results demonstrate that SVD significantly improves cumulative returns across multiple non-stationary environments. Notably, this study represents the first effort to incorporate biologically inspired forgetting into non-stationary reinforcement learning without task labels, achieving both strong performance and inherent adaptability.

Deep Reinforcement LearningEnvironmental DriftForgetting Mechanisms

This work investigates the nature of catastrophic forgetting in continual learning, disentangling the effects of representation loss from interface drift. By splicing the early layers of a pre-trained model with the later layers of a sequentially updated model and introducing task-specific “transfer keys” to align internal interfaces, the method effectively recovers forgotten knowledge. Combining anchor activation pairing with a compact interface alignment operator, the approach demonstrates on ResNet and small Vision Transformers that forgetting primarily stems from interface drift rather than irreversible loss of learned representations, as latent features remain accessible. Experiments on benchmarks such as Split CIFAR-100 show that most of the original performance can be restored, highlighting the critical role of re-indexing latent computations for knowledge recovery.

catastrophic forgettingcontinual learninginterface drift

This work addresses the instability and task-agnostic collapse commonly observed in self-play policy distillation, which often stem from ill-timed updates of the teacher policy. Through a systematic analysis of the temporal coupling between the teacher’s freezing interval (quarantine period) and the student’s learning dynamics, the study identifies clock-driven teacher refreshes as a primary cause of collapse. To mitigate this, the authors propose Consolidation-Gated Teacher Refresh (CGTR), an adaptive gating mechanism that triggers teacher updates only when jointly validated by improvements in reward and safe trajectory length. Requiring no task-specific hyperparameter tuning, CGTR achieves zero collapse across four diverse tasks—Chemistry, Biology, Physics, and ToolUse—while attaining state-of-the-art performance under a unified hyperparameter configuration and automatically adjusting the teacher refresh frequency per task.

isolation periodsself-distillationstate-oblivious collapse

Hot Scholars

IR

Irina Rish

University of Montreal / Mila -Quebec AI Institute
Artificial IntelligenceMachine LearningNeuroscience
EB

Eugene Belilovsky

Associate Professor, Concordia University and Mila Quebec AI Institute
Distributed LearningContinual LearningFederated LearningLearned Optimizers
XZ

Xiaocong Zhao

Tongji University | TU Dresden
autonomous drivinginteractive decision makingdriving behaviour modelling
JG

Jianwei Gong

Beijing Institute of Technology
Intelligent VehicleRobotic VehicleRobot Control
AJ

Ali Jannesari

Associate Professor, Iowa State University
high-performance computingmachine learningparallel computingsoftware analytics