Learning Under Forgetting: Statistical Support-Selective Retention in Stochastic Training Dynamics

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of a unified dynamical explanation for selective retention mechanisms in neural networks by proposing a theoretical framework termed "repeated reinforcement and persistent forgetting." Departing from the conventional view that treats forgetting as a deficiency, this work establishes it as a controllable inductive bias. Grounded in an independent feature model and small-step approximation, a three-tiered theoretical derivation reveals that forgetting induces spectral filtering and minimum description length (MDL)-style compression effects, with conclusions holding irrespective of architecture or scale. Experiments validate the joint selective mechanism of reinforcement and forgetting, demonstrating that forgetting exerts a decisive influence on both the composition of the retained set and its temporal dynamics.
📝 Abstract
Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selective retention remains incomplete. We propose Repeated Reinforcement with Persistent Forgetting (RPF) dynamics, a minimal framework in which repeated exposure reinforces patterns and structures that recur in the data, while persistent forgetting attenuates learned information. This view treats forgetting not merely as a failure mode, but as a selection mechanism. We build the theory in three successive layers. First, in an independent-feature model, we derive an exposure-selective survival law and a support-dependent retention boundary characterizing which patterns persist under forgetting. Second, in a shared-parameter model, we show that forgetting induces spectral filtering over covariance modes, preserving strongly supported shared components while suppressing weak ones. Third, under small-step and norm/coding approximations, we show how RPF dynamics induce an implicit trade-off between data fitting and the cost of stored information, yielding Minimum Description Length (MDL)-like compression. Controlled experiments provide evidence for this reinforcement--forgetting selection mechanism in scalar memories and a nonlinear shared network. Joint reinforcement and attenuation interventions shift conditional retention, while matched exposure counts reveal forgetting-dependent effects of reinforcement timing and changes in the composition of the retained set. Together, these results show that repeated reinforcement and persistent forgetting jointly provide a controllable source of inductive bias beyond neural architecture and scale.
Problem

Research questions and friction points this paper is trying to address.

selective retention
forgetting dynamics
inductive bias
spectral bias
training dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Persistent Forgetting
Repeated Reinforcement
Spectral Filtering
Minimum Description Length
Inductive Bias