🤖 AI Summary
This study addresses the unclear impact of reinforcement learning (RL) on the policy boundaries and coverage of large language model reasoning. Building upon the Qwen and Gemma model families, this work proposes a two-stage autoregressive policy model that decouples strategy selection from execution, complemented by theoretical derivations to elucidate the cross-domain transfer mechanisms of RL. The findings reveal that RL implicitly reweights existing strategies via bias rather than expanding the support set, thereby providing a unified explanation for both performance gains and catastrophic forgetting. Furthermore, this research establishes a theoretical foundation for scaling laws such as log-sigmoid, connecting policy evolution intrinsically to computational efficiency.
📝 Abstract
Recent studies on reinforcement learning (RL) report seemingly conflicting evidence about large language model (LLM) reasoning. Training on mathematics can improve performance in other domains, yet gains in Pass@1 can coincide with lower Pass@$N$ than the base model. This raises a fundamental question: does RL expand an LLM's reasoning boundary, or merely reweight its existing reasoning space? We revisit these phenomena across Qwen and Gemma model families, showing both cross-domain gains and forgetting, while coverage at large sampling budgets increases on some tasks and decreases on others. Detailed analysis of solution traces before and after RL indicates a shift in the reasoning strategies the model employs, motivating a two-stage autoregressive policy model that separates \emph{strategy selection} from problem-specific execution. Within this framework, we prove how RL's implicit bias reshapes strategy preferences, allowing gains on some tasks while suppressing strategies required by others. This mechanism can also broaden or narrow coverage at a given sampling budget even without expanding strategy support. We further provide theoretical justifications for log-sigmoid and log-linear scaling laws in RL compute, and evaluate their predictive power. Together, these results connect changes in strategy selection to cross-domain transfer, reasoning coverage, and compute scaling.