Effective Synthetic Data Curation Requires Group-Level Signals

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of generative capabilities in large language models caused by existing synthetic data training paradigms, which rely on instance-level signals and neglect inter-sample interactions. To overcome this limitation, this work proposes a group-level signal-based framework for evaluating synthetic data utility. By incorporating group-level influence estimation and data composition analysis to capture sample interactions, the method further introduces a low-cost diagnostic technique that prioritizes critical data subsets to optimize selection strategies. This research provides the first empirical evidence demonstrating that group-level signals are essential for effective synthetic data training. Experimental results indicate that the proposed approach significantly enhances downstream task performance and generation quality. Moreover, under constrained computational budgets, it recovers benefits comparable to exhaustive group-level scoring, thereby achieving highly efficient data selection.
📝 Abstract
Synthetic data now is essential to LLM training, used to strengthen advanced capabilities such as autonomous and long-horizon task execution. Yet recent work shows that training on it at scale can degrade model generation, making it important to decide what synthetic data is worth training on. While current data curation practices do so with individual-level signals (i.e., estimates of each data sample's training utility in isolation), across pre-training and post-training settings we show that this is insufficient for synthetic data, and that group-level signals (i.e., estimates of utility that account for interactions among data samples) are necessary for effective data curation. First, we show that individual-level signals are blind to how samples jointly affect training: synthetic datasets with different compositions can be indistinguishable under individual-level influence yet differ sharply under group-level influence, and curating by the latter yields better downstream performance, particularly in generative capability. Second, we find that group-level signals matter more as training pipelines become increasingly synthetic: among widely used data curation methods, only those incorporating them improve over baseline, with gains increasing when weights capturing relations among samples are amplified. Finally, we translate these findings into practice -- for model developers under a compute budget, we offer a cheap diagnostic that prioritizes which groups of synthetic data most need group-level estimation, recovering much of the benefit of full group-level scoring at a fraction of the compute cost.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Data
Data Curation
Large Language Models
Group-Level Signals
Individual-Level Signals
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Curation
Group-Level Signals
Large Language Models
Data Selection
Compute-Efficient Diagnostics
🔎 Similar Papers
No similar papers found.