🤖 AI Summary
This work addresses the prohibitive computational overhead caused by frequent feature recomputation in gradient-based data selection. We propose CDIS, a method grounded in a theoretical proof that cached features preserve ranking order. CDIS reuses stale gradient features and introduces refreshment and affine calibration mechanisms to recover full-computation accuracy at minimal cost. Furthermore, it integrates influence functions, gradient alignment, and source- and length-quota constraints to enable efficient data filtering. Experimental results demonstrate that CDIS achieves a 3.5× speedup over baselines, outperforms random selection by 12 points on GSM8K, and closely approximates the performance of full recomputation, thereby establishing an effective balance between efficiency and accuracy.
📝 Abstract
Gradient-based data selection methods such as LESS score each candidate by the alignment between its gradient and a target validation gradient, and recomputing per-example gradient features at every new checkpoint dominates their cost. Across three selection seeds, two model families, two candidate pools, and two target tasks, features cached at a post-warmup checkpoint and paired with fresh validation gradients preserve the ranking 40 optimizer steps later with Spearman correlation from 0.952 to 0.991, while the top-10% subset they induce misses 10 to 22% of the examples that full recomputation selects. We therefore propose Cached Diverse Influence Selection (CDIS), which recomputes gradient features for the top-ranked fraction $p$ of candidates under the stale scores, fits an affine calibration on the recomputed examples, and selects the final subset under source and length quotas. A refresh fraction at or above the selection fraction recovers the exact top-$k$ subset whenever the calibrated stale scores have bounded error, and the budget curves follow this rule: at $p=0.3$ the recovered top-$k$ subset coincides with full recomputation in every setting, allocating the same budget per stratum recovers the stratified subset at 0.92 to 1.00, and gradient-stage wall-clock drops 3.5 to 3.6 times. Iterating the cache over four checkpoints keeps top-$k$ overlap at 0.98 or higher at 1.9 gradient features per example against 4 for full recomputation, and stale-to-recomputed agreement on the refreshed examples provides a free check for unsafe reuse. Downstream, unconstrained top-$k$ selection collapses to a single data source and scores 13 points below random selection on GSM8K. CDIS scores 12 points above random selection with paired confidence intervals that exclude zero and trails full recomputation by 4.3 points, one training-run standard deviation, at 3.4 times lower selection cost.