Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability and inefficiency in verifier-free reinforcement learning, where probabilistic rewards over increasingly long reasoning chains induce posterior concentration phenomena (PCP). To tackle this issue, the authors provide the first formal definition of PCP and propose the RLCPR framework, which explicitly mitigates the problem through uncertainty-aware sampling and concentration-aware regularization, thereby stabilizing the GRPO training process. Empirical evaluations demonstrate that the proposed approach surpasses state-of-the-art methods by up to 4.0% across six benchmarks while significantly improving token generation efficiency.
📝 Abstract
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
Problem

Research questions and friction points this paper is trying to address.

Verifier-free Reinforcement Learning
Probability-based Rewards
Posterior Concentration Phenomenon
Long-horizon Reasoning
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifier-free Reinforcement Learning
Posterior Concentration Phenomenon
Probability-based Rewards
Uncertainty-aware Sampling
Concentration-aware Regularization
🔎 Similar Papers
No similar papers found.