Adynamical systems view of training generativemodels and the memorization phenomenon

📅 2026-05-19
📈 Citations: 0
Influential: 0
📄 PDF

career value

222K/year
🤖 AI Summary
This work investigates the phenomenon of memory in generative model training—where models persistently output similar samples—and elucidates its underlying mechanism through the lens of dynamical systems theory. By integrating the two-timescale dynamics of stochastic gradient descent (SGD) with structural properties of the loss landscape, the study offers the first unified explanation linking memory effects, double descent, and mode collapse, emphasizing the pivotal role of training dynamics themselves. Building upon Austin’s (2016) loss modeling, Borkar’s (2025a, 2026) theories of collapse and double descent, and recent advances by Azizian et al. (2024) on constant-stepsize SGD, the authors construct a dynamical framework that reveals the fundamental causes of output stagnation during training.
📝 Abstract
Using recent works of one of the authors (VSB) on collapse in generative models and two time scale dynamics in stochastic gradient descent in high dimensions, we give a system theoretic explanation of the memorization phenomenon in generative models. This relies purely on the dynamic aspects of the training phase. Specifically, we use a result of Austin [2016] to motivate a stylized model for the loss function for stochastic gradient descent (SGD) wherein the loss function has a strong dependence on some variables and weak dependence on the rest in a precise sense. This naturally leads to two distinct time scales in the constant step size SGD that is commonly used in machine learning. This fact has been used to explain the double descent phenomenon in SGD in Borkar [2026]. In conjunction with a mathematical model for collapse phenomenon in SGD developed in Borkar [2025a], we analyze the constant step size SGD using the recent results of Azizian et al. [2024] in order to explain the phenomenon of memorization wherein a generative model that is concurrently being tuned yields the same or similar outputs for significant stretches of time. This gives a novel perspective on the aforementioned phenomena reported in machine learning literature and their interrelationships, using a dynamical systems viewpoint.
Problem

Research questions and friction points this paper is trying to address.

memorization phenomenon
generative models
stochastic gradient descent
dynamical systems
two time scale dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

dynamical systems
two time scale dynamics
memorization phenomenon
generative model collapse
stochastic gradient descent