🤖 AI Summary
This study addresses the inefficiency caused by fixed budget allocation during reinforcement learning post-training by proposing CERO, an online dual scheduler that enables dynamic budget pacing throughout the entire training cycle. Methodologically, this work innovatively introduces a concave surrogate utility function to optimize cumulative prompt exposure, and efficiently solves it by integrating Fenchel duality, projected online gradient descent, and a reward-variance feedback mechanism. Experimental results demonstrate that the proposed approach achieves state-of-the-art average performance across multiple mathematical reasoning benchmarks, effectively validating the superiority of adaptive budget scheduling.
📝 Abstract
Adaptive rollout methods for group-relative reinforcement learning typically allocate a fixed per-update budget across prompts. We instead study how to coordinate a finite rollout budget over the entire training horizon. We formulate this problem using a concave surrogate utility of cumulative prompt exposure and introduce CERO, an online primal dual scheduler for prompt admission and budget pacing. In our experiments, each admitted prompt receives a fixed-size response group. CERO instead adapts which prompts are selected, how often they are revisited across rounds, and how many groups are generated in each round. A compact Fenchel representation linearizes the dependence on cumulative exposure, while projected online gradient descent updates prompt-specific supporting slopes and a shared budget price using reward-variation feedback and budget deviations. We establish pathwise guarantees for the surrogate allocation objective against fixed-rate and same-path time-varying benchmarks, with explicit terms for proxy discrepancy and rate variation. Under matched training-response budgets, CERO attains the highest avg@16 macro-average on each of three backbones across five mathematical reasoning benchmarks. Mechanistic analyses link CERO's prompt choices to within-group reward contrast, while multi-seed ablations show gains from adaptive pacing over both uniform and preset spending schedules.