A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops

๐Ÿ“… 2025-02-26
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In self-consuming training loops (STL), models iteratively train on synthetic data, often suffering from performance degradation or catastrophic collapse; however, the underlying collapse mechanisms and differences in robustness remain theoretically unexplained. Method: We introduce the novel concept of *recursive stability* and derive a theoretical bound on STL generalization error. Leveraging Transformer architectural properties and the ratio of real to synthetic data, we conduct rigorous generalization analysis and derive convergence conditions for in-context learning within STL. Contribution: We proveโ€”for the first timeโ€”that a constant fraction of real data suffices to guarantee convergence of Transformers under STL. We identify architecture design and data-mixing strategies as decisive factors for stability. Furthermore, we provide a verifiable collapse criterion and principled guidelines for optimal synthetic-data scaling, establishing a theoretical foundation and practical framework for safe, sustainable self-iterative training.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Adversarial Learning & RobustnessSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labeling
๐Ÿ“ Abstract
High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of recursive stability and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing.
Problem

Research questions and friction points this paper is trying to address.

Prevent model collapse in self-consuming loops
Analyze recursive stability in training loops
Determine optimal real-synthetic data proportion
Innovation

Methods, ideas, or system contributions that make the work stand out.

recursive stability concept
theoretical generalization analysis
optimal synthetic data sizing
๐Ÿ”Ž Similar Papers
No similar papers found.