Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of gradient matching in online data selection, where covariance penalties introduce diagonal noise that compromises sample filtering. To overcome this limitation, we propose LOOM, a framework that eliminates bias at its source via leave-one-out estimation and constructs an unbiased error objective through intra-batch cross-estimation, selecting samples weighted by their signal-to-noise ratio. By integrating an Adam-preconditioned Gram matrix with a greedy algorithm, LOOM enables efficient online filtering. Notably, our approach accurately selects fine-tuning data for large language models without requiring a validation set. Extensive experiments demonstrate that LOOM outperforms both full-data training and existing baselines by 2.4 points across multiple tasks while significantly mitigating the adverse effects of label noise.
📝 Abstract
Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update's error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method{} on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a $(1-e^{-γ})$ guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6--2.0, and selects injected label noise at under a fifth of its base rate.
Problem

Research questions and friction points this paper is trying to address.

online data selection
gradient matching
LLM fine-tuning
covariance penalty
batch selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leave-One-Out Gradient Matching
Online Data Selection
Gradient Gram Matrix
Signal-to-Noise Ratio
LLM Fine-Tuning