🤖 AI Summary
This work investigates the theoretical underpinnings of memorization and overfitting in stochastic interpolation generative models. Focusing on continuous-time stochastic differential equations and their Euler discretization, it provides the first rigorous theoretical definitions of overfitting and underfitting in generative modeling and derives closed-form expressions for the optimal velocity field and score function. The analysis reveals that generated samples can be expressed as training samples perturbed by three controllable error terms, whose bias is jointly determined by the discretization step size and estimation error. Synthetic experiments corroborate the theoretical prediction that generated samples cluster around the training data distribution, highlighting the critical roles of error accumulation and noise modeling in the model’s reconstruction capability.
📝 Abstract
This paper provides a theoretical account of memorization in stochastic interpolation models. By leveraging closed-form expressions for the optimal velocity field and the associated score function, we show that, in the continuous-time oracle setting, both deterministic and stochastic generation processes recover training samples. Under Euler discretization, generated samples remain centered around training samples, with deviations controlled by the step size. We further analyze generation in the presence of estimation errors and show that accumulated estimation errors control the endpoint deviation from the training set. These results imply that the generated sample admits a representation as a training sample perturbed by three controlled terms: a discretization-induced bound, an estimation-error-induced bound, and stochastic Gaussian noise. Based on this characterization, we provide theoretical definitions of overfitting and underfitting in generative models. Synthetic simulations support our theoretical findings.