LESSER: Post-Training Data Selection with Output-Layer Gradients

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational cost of full-parameter gradient-based data selection in large language model post-training, which hinders scalability to large candidate pools. To overcome this limitation, this work proposes LESSER, a method demonstrating that output-layer gradients alone can effectively approximate full-gradient representations for data filtering. Notably, this approximation requires only a forward pass, eliminating the need for backpropagation. Evaluated across supervised fine-tuning (SFT) and reinforcement learning (RL) benchmarks, LESSER reduces the computational overhead of feature extraction by 9.7× and 3.0×, respectively, while maintaining performance comparable to full-gradient approaches. By substantially lowering computational demands without sacrificing effectiveness, this method provides a practical and scalable solution for efficient data selection during the post-training of large language models.
📝 Abstract
The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by $9.7\times$ for SFT and $3.0\times$ for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.
Problem

Research questions and friction points this paper is trying to address.

post-training data selection
gradient-based data selection
large language models
computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Training Data Selection
Output-Layer Gradients
Gradient-based Data Selection
Computational Efficiency
Large Language Models