Strategically Diverse Sampling for Self-Training

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing self-training methods, where repetitive sampling induces homogeneous strategies and over-reliance on inherent model preferences, hindering performance on complex tasks. We propose a self-training principle centered on strategy diversity, introducing GROOT to construct hierarchical method trees that generate diverse reasoning trajectories, combined with Verbalized Sampling to optimize data construction. This work provides the first empirical evidence that strategy diversity is a stronger determinant of self-training efficacy than label correctness or teacher model scale. Experiments demonstrate that our approach significantly outperforms standard i.i.d. training on challenging benchmarks such as competitive programming. Notably, we show that smaller models trained on diversified error trajectories can surpass larger models optimized via conventional knowledge distillation.
📝 Abstract
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
Problem

Research questions and friction points this paper is trying to address.

Self-training
Strategic diversity
Large language models
Sampling
Data construction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Strategic Diversity
Self-Training
GROOT
Hierarchical Tree Sampling
Verbalized Sampling