🤖 AI Summary
This study addresses the challenge of group distributional robustness over multi-source heterogeneous data during LLM post-training, specifically tackling the dynamic minimax regret problem where only immediate mini-batch feedback is available without re-evaluating historical data. We formulate a sampler-optimizer two-player game framework and propose the DUCB-OGD algorithm, which integrates discounted upper confidence bound sampling, online gradient descent, and exponential moving average loss estimation. This work establishes the first theoretical framework for tracking the dynamic worst-case group under stale partial feedback, proving near-optimal regret bounds. Empirical evaluations across supervised fine-tuning, preference optimization, and reinforcement learning tasks demonstrate that our method significantly enhances worst-group robustness with negligible computational overhead while seamlessly integrating into existing training pipelines.
📝 Abstract
Modern LLM training increasingly relies on heterogeneous data sources spanning different domains, tasks, preference distributions, and difficulty levels. We study dynamic minimax regret for group-distributionally robust LLM post-training under instantaneous mini-batch-only bandit feedback. The framework views the training as a two-player sampler-optimizer process: a sampler adaptively selects among data sources using bandit feedback, while an optimizer updates the model parameters using stochastic gradients from the selected source. We focus on the practically restrictive setting where source losses evolve with model training but historical data are not re-evaluated, requiring the sampler to track instantaneous worst-sources from stale partial feedback. We propose DUCB-OGD, a simple and scalable algorithm that couples a Discounted Upper-Confidence-Bound sampler with an Online Gradient Descent optimizer. The sampler maintains exponential moving average loss estimates and confidence radii based on discounted effective sample sizes, avoiding costly re-evaluation of past data or intrusive changes to standard training pipelines. For $K$ data sources and $T$ training steps, we prove that DUCB-OGD achieves a dynamic minimax regret of $\tilde{O}(K^{1/4}T^{3/4})$, which is optimal up to logarithmic factors for the undiscounted objective under our feedback model. Extensive experiments across supervised fine-tuning, preference optimization, and reinforcement learning show that DUCB-OGD integrates seamlessly into modern LLM training pipelines and improves worst-group robustness with negligible computational overhead compared with standard sampling baselines.