All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the policy collapse issue in the late-stage GRPO algorithm during Reinforcement Learning with Verifiable Rewards (RLVR) training for large language models. We propose a unified theoretical framework integrating optimization dynamics and information theory, revealing the mechanism of policy capacity shrinkage and the inherent conflict between policy concentration and task accuracy. Based on these insights, we define the MEI metric as an online early-warning signal and introduce Mesh Learning to expose multiple reasoning paths, preventing any single policy from dominating optimization and establishing policy preservation as a key principle for stable RLVR. Experiments demonstrate that our approach significantly improves performance on benchmarks such as AIME across the Qwen and Phi model families, achieving gains of up to 13.4 percentage points.
📝 Abstract
During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at https://github.com/Ayanami-0123/Open-Mesh-Learning.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning with Verifiable Rewards (RLVR)
Catastrophic Strategy Collapse
Large Language Models
Strategy Capacity
Optimization Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Verifiable Rewards (RLVR)
Catastrophic Strategy Collapse
Mesh Learning
Mirrored Entanglement Index (MEI)
Strategy Capacity
🔎 Similar Papers
No similar papers found.