🤖 AI Summary
This paper investigates the theoretical behavior of generative models under finite training samples. For both deterministic and stochastic generation processes, it derives closed-form solutions for the velocity field and score function within a stochastic interpolation framework—revealing that the former exactly recovers training samples, while the latter corresponds to adding Gaussian noise to them. It introduces the first formal definitions of underfitting and overfitting for generative models, proving that, in the presence of model estimation error, stochastic generation amounts to convex combinations of training samples corrupted by a mixture of noise sources. These theoretical findings are empirically validated on downstream classification tasks, confirming that the characterized noise structure aligns with observed generalization performance. The core contribution is an analytical theory of generative processes under finite-sample regimes, unifying the explanatory frameworks for sample recovery and perturbation, and providing verifiable criteria for diagnosing underfitting and overfitting.
📝 Abstract
This paper investigates the theoretical behavior of generative models under finite training populations. Within the stochastic interpolation generative framework, we derive closed-form expressions for the optimal velocity field and score function when only a finite number of training samples are available. We demonstrate that, under some regularity conditions, the deterministic generative process exactly recovers the training samples, while the stochastic generative process manifests as training samples with added Gaussian noise. Beyond the idealized setting, we consider model estimation errors and introduce formal definitions of underfitting and overfitting specific to generative models. Our theoretical analysis reveals that, in the presence of estimation errors, the stochastic generation process effectively produces convex combinations of training samples corrupted by a mixture of uniform and Gaussian noise. Experiments on generation tasks and downstream tasks such as classification support our theory.