🤖 AI Summary
This study addresses the issue in reinforcement learning where group-relative objectives bias learning signals toward high-frequency patterns, thereby limiting reasoning diversity. To overcome this, we propose ExPPO, a method that redistributes advantage credit based on the surprisal of prompt-conditioned responses. By introducing a lightweight advantage reshaping rule combined with bounded shaping and shared normalization mechanisms, ExPPO optimizes exploration strategies while preserving verifier polarity. Furthermore, we theoretically derive local conditions for entropy increase and the discovery of correct patterns. Experimental results demonstrate that ExPPO significantly improves both in-domain and out-of-domain reasoning coverage, aggregated accuracy, and diversity under large sampling budgets, effectively increasing the generation of correct patterns.
📝 Abstract
Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using prompt-relative, length-normalized response surprisal and prompt pass rate. ExPPO combines bounded shaping with shared normalization to preserve verifier polarity and approximately maintain each prompt group's total absolute sequence-advantage mass. Our analysis characterizes response-level credit allocation alongside sampled mode updates, deriving local conditions for gains in entropy and correct-mode discovery. Experiments show improved in-domain and out-of-domain reasoning coverage, higher aggregate response accuracy, and strong coverage at large sampling budgets. A controlled multi-answer evaluation further demonstrates increased correct-mode yield and gains in diversity among verified-correct responses. Code is available at https://github.com/jinhangzhan/ExPPO