The Model Knows Another Way: Strategy Switching for Effective RLVR Exploration

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient exploration in Reinforcement Learning with Verifiable Rewards (RLVR) caused by sampling bias, which induces a “dead saturation” phenomenon that hinders the elicitation of alternative reasoning strategies already acquired by the model. To overcome this limitation, we propose Policy-Switching as Reasoning Arms (PSRA), which pioneers formulating strategy switching as competing exploration arms and leverages Bayesian sequential decision-making to dynamically optimize the allocation of exploratory resources between unguided and strategy-conditioned prompting. Our findings reveal that limited sampling obscures the latent capabilities of models. Extensive experiments on the Qwen2.5 model series demonstrate that PSRA effectively mitigates training stagnation, yielding significant improvements in reasoning performance, out-of-distribution generalization, and returns under high computational budgets.
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) is often limited by insufficient exploration: difficult problems can yield uniformly incorrect rollout groups and therefore little learning signal. We show that such failures need not reflect missing capability. Instead, finite sampling often concentrates on a problem-specific dominant reasoning strategy while leaving alternative strategies already supported by the model unexplored. Moreover, the accessibility of these strategies evolves during RL: some are internalized into autonomous behavior, while others become difficult to elicit before being absorbed. Motivated by these observations, we introduce Problem--Strategy Rollout Allocation (PSRA), which treats unguided and strategy-conditioned prompts as competing exploration arms and uses Bayesian sequential allocation to direct a fixed rollout budget toward arms most likely to yield informative, non-saturated groups. A preservation objective keeps useful strategy-conditioned routes accessible while successful guided behaviors are transferred to the unguided policy. Across Qwen2.5 models from 1.5B to 7B and two RL training corpora, PSRA consistently improves reasoning performance, reduces dead saturation, strengthens out-of-distribution transfer, and maintains larger gains under increased inference budgets.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards (RLVR)
Insufficient Exploration
Strategy Switching
Reasoning
Dead Saturation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards
Strategy Switching
Bayesian Sequential Allocation
Problem-Strategy Rollout Allocation
Exploration
🔎 Similar Papers
2024-07-09Neural Information Processing SystemsCitations: 3
Jin Cui
Jin Cui
Principal Engineer
Embedded SystemOS Kernel & DriverHypervisor & VirtualizationComputer uArch modellingFPGA & EDA
X
Xinyue Long
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; School of Software Engineering, Xi’an Jiaotong University
B
Boran Zhao
State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University; School of Software Engineering, Xi’an Jiaotong University
Pengju Ren
Pengju Ren
Professor, Xi'an Jiaotong University
H
Hao Dong
ELLIS Institute Finland and Tampere University