🤖 AI Summary
This study addresses the stability and diversity challenges of networked generative models within self-consuming training loops. Moving beyond the limitations of isolated models, this work introduces the first directed weighted graph-theoretic framework that integrates graph modeling with dynamical systems analysis to characterize the long-term evolutionary dynamics of synthetic data flow and iterative retraining across multiple models. We rigorously establish sufficient conditions for system convergence and reveal how network interaction topologies, the proportion of real data access, and cross-model consumption patterns jointly shape system stability and diversity. By elucidating these underlying mechanisms, this research provides a theoretical foundation for understanding and governing multi-model self-consuming ecosystems.
📝 Abstract
The widespread deployment of generative AI has made it increasingly difficult to distinguish synthetic content from real data. Consequently, synthetic data is inevitably incorporated into the training pipelines of future model generations, forming a self-consuming training loop. Prior work has studied the effects of such recursive self-consuming training, but analyses have largely been limited to isolated models, where a model consumes only its own synthetic data, or to simplified interactions between two models. This paper takes a first step toward understanding networked self-consuming generative models, in which multiple models consume synthetic data generated by one another through complex interaction pathways. We introduce a theoretical framework representing models as nodes in a directed, weighted graph, with edge weights governing the flow of synthetic data among models. Using this framework, we analyze the long-term behavior of networked models under retraining dynamics, establishing conditions for convergence and characterizing the resulting fixed points. We further investigate how the system's long-term stability and diversity are shaped by each model's access to real data, cross-model data consumption, and the structure of the interaction graph.