Limits of Difficulty Scaling: Hard Samples Yield Diminishing Returns in GRPO-Tuned SLMs

📅 2026-04-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the effectiveness of preference optimization for small language models (≤3B parameters) on challenging mathematical reasoning tasks under resource constraints. Leveraging Group Relative Policy Optimization (GRPO) with LoRA fine-tuning, the authors conduct difficulty-stratified training and evaluation on the GSM8K and MATH datasets. The findings reveal that performance gains on high-difficulty samples exhibit diminishing returns, and training exclusively on low-difficulty data—using only about 45% of the total training steps—achieves comparable accuracy to full-dataset training. Moreover, models trained on GSM8K demonstrate strong cross-dataset generalization, outperforming baselines by 3–5% on the numerical subset of MATH, highlighting the efficacy of targeted, difficulty-aware optimization for resource-efficient alignment in mathematical reasoning.

Technology Category

Machine Learning: OptimizationNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this claim in a resource-constrained setting by applying GRPO with LoRA to SLMs (up to 3B) for math reasoning on GSM8K and MATH datasets with difficulty-stratified analyses. As problem difficulty increases, accuracy plateaus, revealing a capacity boundary: GRPO primarily reshapes output preferences without reliably improving hardest-tier solving. Consistent with this, training GRPO only on lower-difficulty problems matches full-dataset accuracy across difficulty tiers while using only ~45% training steps, indicating diminishing returns from harder samples in this regime. We also find a cross-dataset generalization effect: GSM8K-trained GRPO achieves higher accuracy on the numeric subset of MATH than MATH-trained GRPO, exceeding it by ~5% at 1.5B and by ~3% at 3B. We show that the best achievable gains depend strongly on the base model's prior reasoning competence and the dataset's difficulty profile.
Problem

Research questions and friction points this paper is trying to address.

difficulty scaling
preference optimization
small language models
diminishing returns
math reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

GRPO
difficulty scaling
diminishing returns
small language models
cross-dataset generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Suraj Yadav
IIIT Delhi
Siddharth Yadav
Siddharth Yadav
IIIT Delhi
Deep LearningSystems
P
Parth Goyal
IIIT Delhi