When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling

๐Ÿ“… 2026-07-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the frequent conflation of source expansion (SE) and fixed-source synthesis (FSS) in existing literature, which has hindered a systematic understanding of FSS mechanisms. To isolate FSS, we fix both the seed problem pool and the teacher model, varying only the per-problem generation budget. Through a combination of rejection sampling, controlled ablation experiments, and comparative synthesis protocols, we systematically investigate FSS scaling behavior. We propose a revised scaling law tailored to FSS and validate its predictive accuracy across diverse teacherโ€“student model pairs. Our findings reveal that FSS matches SE performance under low budgets but is outperformed by SE at higher budgets. Moreover, within FSS, simple rejection sampling consistently surpasses more complex strategies, suggesting an inherent performance ceiling intrinsic to the FSS paradigm.
๐Ÿ“ Abstract
Synthetic data can be scaled along two routes: Source Expansion (SE), which enlarges the source by adding seed materials or generators, and Fixed-Source Synthesis (FSS), which holds the source fixed and scales the generation budget. Existing scaling studies typically expand the source as the data grows, conflating SE with FSS and leaving FSS underexplored. We isolate FSS by holding the seed-question pool and teacher model fixed, varying only the per-question response budget under Rejection Sampling (RS). We adapt the rectified scaling law to FSS, deriving it from how repeated sampling covers a fixed source. Empirically, the derived form, fit on low budgets, predicts performance at the held-out highest budget for every evaluated teacher--student pair. At matched total-sample budgets, SE and FSS are comparable at small budgets; at large budgets, adding seed questions outperforms spending the same budget on more responses. Within FSS, however, neither synthesizing additional questions from the existing seeds nor varying the synthesis protocol outperforms plain RS at matched budgets. FSS is thus a bounded scaling axis and a controlled setting for comparing synthesis protocols. We will release our code and data to facilitate further research.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Data
Fixed-Source Synthesis
Source Expansion
Scaling Law
Rejection Sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fixed-Source Synthesis
Source Expansion
Synthetic Data Scaling
Rejection Sampling
Scaling Laws
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
X
Xu Guo
Shanghai AI Laboratory; Shanghai Innovation Institute; Fudan University
J
Jian Tong
Shanghai AI Laboratory
Z
Zhihui Lu
Fudan University
Qipeng Guo
Qipeng Guo
Fudan University