Does Scaling Reinforcement Learning Really Require More Training?

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of conventional inference scaling, which relies on increased computation and faces inherent performance ceilings under fixed reinforcement learning (RL) histories. To overcome this, it introduces the concept of policy space expansion and proposes the SURGE framework, which uniquely treats completed RL training histories as reusable resources. By employing spectral decomposition to extract dominant anchor components and complementary donor components, and by adaptively determining block sizes based on weights, SURGE achieves gradient-free checkpoint merging that enhances model capabilities without additional training. The proposed method significantly outperforms input checkpoints on mathematical and coding benchmarks—achieving 54.17% on DeepSeek AIME24—while consuming fewer inference tokens. This work establishes a novel paradigm for inference scaling that requires no extra computational overhead.
📝 Abstract
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dominant component and incorporate the donor's complementary component. With a fixed target for how much of the anchor update to retain, SURGE determines the block size from the weights without testing candidate policies. We evaluate two 1.5B mathematical-reasoning histories, DeepSeek and Nemotron, and one 7B coding history, OLMo. SURGE improves benchmark-average accuracy over both input checkpoints while using fewer reasoning tokens than the anchor. It reaches 54.17% on DeepSeek AIME24 against a measured native maximum of 50.83%, and 83.7% on OLMo HumanEval+ against 82.8%. These gains exceed the observed training curves. Geometric controls support the importance of RL-update structure beyond weight displacement or token reduction alone. Each constructed model runs as a single policy. Our findings identify stored RL history as a reusable scaling resource: the capability available from a training run need not end at its best checkpoint.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Scaling Reasoning
Policy-space Scaling
Checkpoint Fusion
Mathematical Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy-space scaling
SURGE
Eigenspace fusion
Gradient-free RL scaling
Checkpoint merging
🔎 Similar Papers