🤖 AI Summary
This study investigates whether language models undergo self-reinforcing cycles of social bias when continually pretrained on their own synthetically generated data. Leveraging the Bias in Bios dataset, we design a controlled iterative training pipeline that integrates synthetic data generation, quantitative bias evaluation, and comparison against standard language modeling metrics. Our analysis reveals, for the first time, a phenomenon we term “fairness collapse”: social biases are silently amplified across training iterations well before any noticeable degradation in overall model performance. These findings demonstrate that contamination by synthetic data can insidiously exacerbate societal biases, underscoring the inadequacy of conventional language modeling metrics alone in detecting fairness-related degradation.
📝 Abstract
Generative models trained on artificially generated data have been shown to exhibit model collapse, resulting in significant performance degradation. As synthetic content increasingly contaminates the training corpora of language models, this raises critical concerns about the use of open data in continued pretraining. Although previous work has demonstrated model collapse in language models, it remains unclear whether exposure to synthetic data amplifies or attenuates the social biases already present in pretrained models. Because language models are known to reproduce and amplify demographic stereotypes, recursive training on self-generated data may create a self-reinforcing feedback loop in which biased associations become progressively stronger across generations. We call this hypothesized phenomenon fairness collapse. In this work, we construct controlled training regimes in which models are repeatedly trained on synthetic data using the Bias in Bios dataset. Across experiments, we observe a consistent and concerning pattern: fairness degradation emerges before substantial degradation is reflected by standard language-modeling metrics. This result highlights a critical risk associated with synthetic data contamination in language model training: bias can increase silently before strong indicators of model collapse become apparent.