How to Synthesize Text Data without Model Collapse?

📅 2024-12-19
🏛️ arXiv.org
📈 Citations: 6
✨ Influential: 2
📄 PDF
🤖 AI Summary
This work addresses “model collapse”—the progressive degradation in performance observed when language models are iteratively trained on synthetic text data—by identifying a negative correlation between synthetic data proportion and model performance, alongside n-gram over-concentration. We propose a token-level editing method grounded in human-written text and provide the first theoretical proof that this strategy strictly bounds test error, thereby provably preventing collapse. Furthermore, we introduce a distribution-aware semi-synthetic data paradigm that overcomes the inherent degeneration bottleneck of purely generative data. Through multi-stage pretraining and fine-tuning experiments across diverse downstream tasks, our approach significantly mitigates model collapse, yielding up to a 3.2% absolute accuracy improvement. Empirical results demonstrate the superiority and robustness of semi-synthetic data over fully synthetic alternatives.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Diffusion Models for Vision

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-${n}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance.
Problem

Research questions and friction points this paper is trying to address.

Impact of synthetic data on language model training
Preventing model collapse in synthetic data synthesis
Token-level editing to improve model performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token editing prevents model collapse
Semi-synthetic data balances distribution
Finite upper bound ensures stability
🔎 Similar Papers
No similar papers found.
Shanghai Jiao Tong University | BIGAI | Peking University | Tsinghua University | Shanghai Artificial Intelligence Laboratory
Xuekai Zhu
Xuekai Zhu
Shanghai Jiao Tong University
Synthetic DataReasoningLanguage Model
Daixuan Cheng
Daixuan Cheng
Gaoling School of AI, Renmin University of China
LLM Pre-TrainingDomain AdaptationReasoning
Hengli Li
Hengli Li
Institute for Artificial Intelligence, Peking University
Machine LearningNatural Language Processing
Kaiyan Zhang
Kaiyan Zhang
Tsinghua University
Foundation ModelCollective IntelligenceScientific Intelligence
Ermo Hua
Ermo Hua
Tsinghua University
Physics-driven Foundation Model
Xingtai Lv
Xingtai Lv
Tsinghua University
Large Language ModelNatural Language Processing
N
Ning Ding
Department of Electronic Engineering, Tsinghua University; Shanghai Artificial Intelligence Laboratory
Z
Zhouhan Lin
LUMIA Lab, Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory
Z
Zilong Zheng
State Key Laboratory of General Artificial Intelligence, BIGAI
B
Bowen Zhou
Department of Electronic Engineering, Tsinghua University; Shanghai Artificial Intelligence Laboratory