CRAFT: Clustered Regression for Adaptive Filtering of Training data

📅 2026-04-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently selecting high-quality subsets from ultra-large corpora to reduce fine-tuning costs. The authors propose CRAFT, a method that decomposes the source–target joint distribution and employs a two-stage strategy: first allocating the selection budget across clusters according to k-means proportions to match the source distribution of the validation set, then selecting within each cluster samples whose target embeddings best approximate the validation target distribution. CRAFT is the first approach to combine clustering with conditional expectation distance for data selection. Theoretically, proportional allocation bounds the KL divergence regardless of the specific embedding scheme. Evaluated on fine-tuning mBART with 33 million sentence pairs, CRAFT achieves a BLEU score of 43.34—outperforming TSDS by 2.13 points—and runs 2.8× faster than TAROT, completing the entire pipeline on CPU in under one minute.

Technology Category

Machine Learning: Mixture of Experts (MoE)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageConstraint Satisfaction and Optimization: Distributed CSP/Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Selecting a small, high-quality subset from a large corpus for fine-tuning is increasingly important as corpora grow to tens of millions of datapoints, making full fine-tuning expensive and often unnecessary. We propose CRAFT (Clustered Regression for Adaptive Filtering of Training data), a vectorization-agnostic selection method for training sequence-to-sequence models. CRAFT decomposes the joint source-target distribution and performs a two-stage selection: (i) match the validation source distribution through proportional budget allocation across k-means clusters, and (ii) within each source cluster, select training pairs whose target embeddings minimize a conditional expected distance derived from the validation target distribution. We prove that proportional cluster allocation bounds the continuous KL divergence between selected and validation distributions, with the residual controlled by cluster diameters. We evaluate CRAFT on English-Hindi translation by selecting training data from 33 million NLLB sentence pairs and fine-tuning mBART via LoRA. CRAFT achieves 43.34 BLEU, outperforming TSDS (41.21) by 2.13 points on the same candidate pool and encoder while completing selection over 40 times faster. With TF-IDF vectorization, the entire pipeline completes in under one minute on CPU. TAROT achieves 45.61 BLEU, but CRAFT completes selection in 26.86 seconds versus TAROT's 75.6 seconds, a 2.8 time speedup.
Problem

Research questions and friction points this paper is trying to address.

data selection
fine-tuning
training data filtering
large-scale corpora
subset selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

CRAFT
training data selection
clustered regression
adaptive filtering
vectorization-agnostic
🔎 Similar Papers
No similar papers found.
P
Parthasarathi Panda
Google, USA
A
Asheswari Swain
Department of Computer Science & Information Systems, BITS Pilani, Hyderabad, India
S
Subhrakanta Panda
Department of Computer Science & Information Systems, BITS Pilani, Hyderabad, India