🤖 AI Summary
This study addresses the diminishing performance gains in large language model (LLM) training caused by redundancy and errors in synthetic data. It establishes, for the first time, a linear theoretical framework from a training dynamics perspective to characterize the value of synthetic data, explicitly defining optimal addition quantities and marginal utility. Based on this theory, we propose Training-Aware Target Coverage (TATC), a method that balances input coverage with error conditions to precisely select high-quality samples capable of effectively expanding directional coverage over target tasks for fine-tuning. Experiments demonstrate the validity of our theoretical analysis and show that TATC significantly outperforms existing baselines on the GSM8K mathematical reasoning task, substantially enhancing the performance of the Qwen2.5-Math model.
📝 Abstract
Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.