Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that noise and distribution shifts in synthetic data degrade training utility, while existing methods overlook the dynamic needs of the learner. To this end, we propose FROST, a framework that evaluates synthetic sample utility online by anchoring to real-data gradients. Without requiring external validators or held-out sets, FROST leverages real-time gradient signals to calibrate batch utility and adaptively filter low-quality samples. Empirical evaluations demonstrate that filtering merely 20%–30% of the data yields substantial performance gains on image classification and Text-to-SQL tasks. Furthermore, when deployed in an industrial advertisement re-ranking system, FROST significantly outperforms the production baseline, highlighting its practical efficacy.
📝 Abstract
Synthetic data can scale training supervision when real-world data are limited, but noise and distribution mismatch can reduce its value. Existing synthetic data selection methods often emphasize fidelity or diversity rather than the learner's evolving needs. We propose FROST, an online framework that estimates synthetic-data utility through gradient feedback anchored in real training data. It calibrates batch utility against recent history to determine when filtering is needed and filters samples only in out-of-band batches to determine what to retain, without an external verifier or held-out validation set. Experiments on two public benchmarks for image classification and LLM fine-tuning for text-to-SQL show that FROST filters out around 20--30% of the synthetic data while improving real-task performance compared with training on the full synthetic data pool. We further apply FROST during training in a large-scale industrial ads re-ranking system, achieving significant performance gains over a highly optimized production baseline, demonstrating its effectiveness and generalizability.
Problem

Research questions and friction points this paper is trying to address.

Synthetic Data
Data Filtering
Data Selection
Distribution Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Filtering
Online Utility Estimation
Gradient Feedback
Real-Anchored Calibration
Data Selection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yanran Wu
Purdue University
S
Sana Lakdawala
Meta
R
Renzo Tassara Miller
Meta
C
Chongyang Bai
Meta
S
Sharath Ciddu
Meta
S
Shivendra Pratap Singh
Meta
K
Kungang Li
Meta
Sandeep Pandey
Sandeep Pandey
TU Ilmenau
Applied Machine LearningComputational Fluid DynamicsDirect Numerical SimulationHigh Performance Computing
Chunwei Liu
Chunwei Liu
Massachusetts Institute of Technology
DatabasesCompound AI SystemsLLMData CompressionIoT