🤖 AI Summary
This study addresses the insufficient exploration in Reinforcement Learning with Verifiable Rewards (RLVR) caused by sampling bias, which induces a “dead saturation” phenomenon that hinders the elicitation of alternative reasoning strategies already acquired by the model. To overcome this limitation, we propose Policy-Switching as Reasoning Arms (PSRA), which pioneers formulating strategy switching as competing exploration arms and leverages Bayesian sequential decision-making to dynamically optimize the allocation of exploratory resources between unguided and strategy-conditioned prompting. Our findings reveal that limited sampling obscures the latent capabilities of models. Extensive experiments on the Qwen2.5 model series demonstrate that PSRA effectively mitigates training stagnation, yielding significant improvements in reasoning performance, out-of-distribution generalization, and returns under high computational budgets.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.