From Collapse to Improvement: Statistical Perspectives on the Evolutionary Dynamics of Iterative Training on Contaminated Sources

📅 2026-02-11
📈 Citations: 0
Influential: 0
📄 PDF

career value

202K/year
🤖 AI Summary
This study investigates the evolutionary dynamics of generative models trained iteratively on synthetic data contaminated with real data, aiming to mitigate model collapse induced by data pollution. Through statistical modeling, mixture distribution analysis, and theoretical analysis of iterative training dynamics—complemented by theoretical derivations and simulations based on next-token prediction language models—the work demonstrates that model collapse can be effectively avoided and the true data distribution even recovered, provided the mixture weight of real data remains non-zero over time and is paired with sufficient sample sizes. This mechanism consistently enhances performance across diverse model classes, offering both theoretical guarantees and practical guidance for sustainable iterative training.

Technology Category

Application Category

📝 Abstract
The problem of model collapse has presented new challenges in iterative training of generative models, where such training with synthetic data leads to an overall degradation of performance. This paper looks at the problem from a statistical viewpoint, illustrating that one can actually hope for improvement when models are trained on data contaminated with synthetic samples, as long as there is some amount of fresh information from the true target distribution. In particular, we consider iterative training on samples sourced from a mixture of the true target and synthetic distributions. We analyze the entire iterative evolution in a next-token prediction language model, capturing how the interplay between the mixture weights and the sample size controls the overall long-term performance. With non-trivial mixture weight of the true distribution, even if it decays over time, simply training the model in a contamination-agnostic manner with appropriate sample sizes can avoid collapse and even recover the true target distribution under certain conditions. Simulation studies support our findings and also show that such behavior is more general for other classes of models.
Problem

Research questions and friction points this paper is trying to address.

model collapse
iterative training
synthetic data
contaminated sources
generative models
Innovation

Methods, ideas, or system contributions that make the work stand out.

model collapse
iterative training
synthetic data contamination
statistical perspective
mixture distribution