PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing data selection methods predominantly rely on fixed heuristics, which struggle to adapt across diverse downstream tasks and often overlook the fundamental differences in learning objectives between language modeling and reasoning tasks. To address these limitations, this work proposes a task-aware and budget-aware perplexity (PPL)-based data selection framework that, for the first time, integrates both task type and data budget into the PPL scoring mechanism. This approach dynamically evaluates sample difficulty, enabling efficient and interpretable data filtering. Empirical results demonstrate that the method achieves state-of-the-art performance on GSM8K using only 1% of the training data and surpasses full-data fine-tuning by 0.9 and 4.8 points on GSM8K and MATH, respectively, when trained on just 10% of the data.
📝 Abstract
Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typically score the entire training sequence and ignore the difference in learning objectives of language modeling and reasoning tasks. In this paper, we propose PPL-Factory, a simple and interpretable data selection framework that combines task-aware perplexity-based scores and data budget-aware selection criteria. Experiments on GSM8K demonstrate that PPL-Factory outperforms other state-of-the-art data selection methods using only $1\%$ of the training set. With $10\%$ of the data, PPL-Factory exceeds full-data fine-tuning accuracy by 0.9 on GSM8K and 4.8 on MATH. Overall, our results demonstrate that task-aware and budget-aware perplexity-based selection provides an effective and applicable approach for efficient fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

data selection
large language models
fine-tuning
perplexity
task-aware
Innovation

Methods, ideas, or system contributions that make the work stand out.

task-aware
budget-aware
perplexity-based selection
data selection
efficient fine-tuning
🔎 Similar Papers