🤖 AI Summary
Large language models face significant challenges in multi-step reasoning due to reliance on sparse terminal rewards, which leads to credit assignment difficulties, high gradient variance, and unstable training dynamics. This work proposes Implicit Behavioral Policy Optimization (IBPO), a novel framework that introduces, for the first time, a counterfactual trajectory comparison mechanism. By sampling multiple reasoning trajectories from the same input and leveraging their differences, IBPO implicitly estimates step-level advantages, thereby transforming sparse terminal rewards into fine-grained, step-sensitive learning signals. This approach effectively reduces gradient variance and substantially enhances both training stability and performance ceilings. Experimental results demonstrate that IBPO significantly outperforms existing methods on mathematical and code reasoning benchmarks, while also exhibiting superior continual learning capabilities and a higher upper bound on reasoning performance.
📝 Abstract
In sparse termination rewards, intra-group comparisons have become the dominant paradigm for fine-tuning reasoning models via reinforcement learning. However, long-term training often leads to issues like ineffective update accumulation (learning tax), solution probability drift, and entropy collapse. This paper presents a necessary condition for algorithm design from a token-level credit assignment perspective: to prevent reward-irrelevant drift, intra-group objectives must maintain gradient exchangeability across token updates, enabling gradient cancellation on weak-credit/high-frequency tokens. We show that two common mechanisms disrupting exchangeability make"non-cancellation"a structural norm. Based on this, we propose minimal intra-group transformations to restore or approximate the cancellation structure in the shared token space. Experimental results demonstrate that these transformations stabilize training, improve sample efficiency, and enhance final performance, validating the value of this design condition.