OptiSelect: How does the Optimizer Shape Data Curriculum?

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight in online data selection for large model pretraining, where the optimizer's reshaping effect on gradients is typically ignored. We propose OptiSelect, a paradigm that establishes an optimizer-aware utility scoring theory to systematically reveal how optimizers shape data curricula. Our analysis uncovers, for the first time, discriminative collapse in sign-based and extreme tangent preconditioners such as Lion and Muon, while theoretically proving that diagonal adaptive optimizers like AdamW yield superior geometric upper bounds for scoring. Experiments on 124M- and 720M-parameter LLMs validate these theoretical derivations, confirming AdamW as the optimal scoring geometry. Furthermore, it maintains significant training efficiency gains even under data restatement scenarios.
📝 Abstract
Online data selection has demonstrated substantial efficiency gains for LLM pretraining by training on the most valuable candidates within each batch. Since a candidate's value is realized through its effective model update, principled selection should account for the optimizer step, which reshapes the raw gradient before it updates model parameters. We formalize this optimizer-aware selection paradigm as OptiSelect and present the first systematic study of how the optimizer shapes data selection. Our theory establishes a selection gain principle in which the advantage of online selection is governed by the discriminability of the optimizer-induced utility scores. We prove that sign-based and polar-tangential preconditioners of Lion and Muon would suffer from a discriminability collapse which caps attainable gains from OptiSelect, whereas diagonal-adaptive optimizers such as AdamW and Sophia admit strictly better upper bounds. The proposed principle also yields a quantitative derivation of the optimal candidate oversampling ratio. Pretraining experiments on 124M and 720M models are consistent with our theoretical analysis and show that AdamW's diagonal-adaptive scoring geometry remains the strongest scoring geometry even with Muon as optimizer. We further demonstrate that OptiSelect retains its benefits under data rephrasing, a technique used in modern data processing pipelines. Our findings provide theoretical foundations and practical guidance for co-designing optimizers and data selection in LLM pretraining.
Problem

Research questions and friction points this paper is trying to address.

Online data selection
Optimizer
LLM pretraining
Data curriculum
OptiSelect
Innovation

Methods, ideas, or system contributions that make the work stand out.

Online Data Selection
Optimizer-aware
Discriminability Collapse
LLM Pretraining
Selection Gain Principle