SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the gradient attenuation and contribution distortion caused by uniform sampling for low-success-rate prompts in maximum likelihood reinforcement learning. To mitigate this, it proposes a scale-balanced sampling allocation strategy that redistributes a fixed budget to equilibrate finite-sample scaling factors across prompts. A minimax allocation model is formulated, yielding an analytical water-filling solution, which is further refined through continuous relaxation theory and multiplicity correction techniques to enhance gradient estimation accuracy. Experimental evaluations on ImageNet, maze navigation, and mathematical reasoning tasks demonstrate that the proposed approach significantly improves gradient alignment and multi-solution coverage.
📝 Abstract
Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.
Problem

Research questions and friction points this paper is trying to address.

Maximum Likelihood Reinforcement Learning
finite rollout budget
gradient attenuation
rollout allocation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Maximum Likelihood Reinforcement Learning
Scale-Equalized Rollout Allocation
Waterline Solution
Multiplicity Correction
Max-Min Optimization
🔎 Similar Papers
No similar papers found.