adaptive credit policy optimization

Designs and implements policy-optimization algorithms that adaptively assign credit across decisions and timesteps by modulating update weights based on rollout outcomes, model confidence, and uncertainty. Builds mechanisms to downweight overconfident decisions and emphasize uncertain decisions that contributed to successful rollouts while preserving the leading-order policy-gradient direction.

adaptivecreditpolicyoptimization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of credit assignment in long-horizon language agent tasks under sparse and delayed rewards, where existing dense credit methods suffer from instability due to interference from infrequent, high-impact actions. To mitigate this, the paper proposes Evidence-Calibrated Policy Optimization (ECPO), a critic-free policy optimization algorithm that statistically calibrates step-level credit prior to policy updates. ECPO employs action-group-based credit estimation contraction and variance-aware anchor weighting to suppress bias from low-frequency actions and noise in anchor selection. This approach significantly enhances training stability and task performance. Evaluated on ALFWorld and WebShop using Qwen2.5-1.5B, ECPO outperforms strong baselines such as GiGPO by 5.2 and 7.3 percentage points in success rate, respectively, with only a 0.1% increase in computational overhead.

credit assignmentLLM agentslong-horizon reinforcement learning

Policy Optimization Algorithms in a Unified Framework

Apr 04, 2025
SW
Shuang Wu
🏛️ Huawei

Policy optimization algorithms suffer from poor interpretability and error-prone implementation due to the complexity of Markov decision process (MDP) modeling and inconsistent use of discounted versus average-reward settings. Method: This paper introduces a unified analytical framework that, for the first time, systematically integrates generalized ergodicity theory with perturbation analysis to characterize the steady-state behavior of diverse policy optimization algorithms under both discounted and average-reward criteria. Contribution/Results: The framework clarifies fundamental algorithmic principles, identifies and corrects common implementation pitfalls, and significantly enhances interpretability and robustness. Empirical validation on MDP modeling and linear quadratic regulator (LQR) benchmarks confirms the framework’s ability to capture algorithmic consistency. Quantitative analysis further demonstrates that minor adjustments to key design parameters exert decisive influence on convergence properties and performance.

Clarify policy optimization algorithms' complex calculations and setupsReduce misuse and improve accessibility of optimization algorithmsUnify framework using ergodicity theory and perturbation analysis

This work addresses the token-level credit assignment challenge in large language models trained with reinforcement learning under sparse rewards. To this end, the authors propose an adaptive credit assignment framework based on fine-grained surrogate entropy. The method employs asymmetric policy modulation—enhancing exploration of uncertain tokens in successful trajectories while suppressing overconfident tokens in failed ones—to achieve precise credit allocation. Additionally, it incorporates modality alignment constraints and proximal policy updates to stabilize gradient directions and mitigate non-local interference. Experimental results demonstrate that the proposed approach significantly outperforms strong baselines such as DAPO, GTPO, and SAPO on mathematical and code generation benchmarks, including AIME 2025 and HumanEvalPro.

credit assignmententropy optimizationlarge language models

This work addresses the limitations of existing reinforcement learning approaches in long-chain reasoning tasks, where coarse-grained, sequence-level credit assignment hinders the identification of critical reasoning steps, and standard KL-divergence penalties often lead to gradient instability and overly conservative policies. To overcome these challenges, the paper proposes a novel critic-free reinforcement learning framework that reframes distributional deviation not as a rigid penalty but as a guiding signal, enabling fine-grained, step-level credit assignment. By eliminating conventional KL constraints, the method effectively mitigates gradient instability, promotes policy diversity, and substantially enhances both the identification of pivotal reasoning steps and overall performance on complex reasoning tasks.

Chain of Thoughtcredit assignmentKL divergence

This work addresses the challenges of uneven rollout allocation and dynamic imbalance in policy optimization that hinder reinforcement learning for reasoning tasks—such as uniform rollouts disregarding gradient variance disparities, softmax-induced gradient attenuation for high-confidence actions, and training instability. To tackle these issues, the authors propose DynaMO, a framework that dynamically allocates rollouts at the sequence level by minimizing gradient variance, introduces advantage modulation at the token level to compensate for gradient decay, and stabilizes update magnitudes through entropy variation monitoring. The key contributions include the first theoretical derivation of a Bernoulli-variance-based rollout allocation criterion and a novel gradient-aware advantage modulation mechanism. Experiments demonstrate that DynaMO significantly outperforms existing RLVR methods across multiple mathematical reasoning benchmarks, achieving both high efficiency and robustness.

gradient attenuationgradient variancepolicy optimization

Latest Papers

What's happening recently
View more

This work addresses the challenge in reinforcement learning of disentangling an agent’s policy contributions from environmental stochasticity in credit assignment. To this end, the authors propose a novel causal inference–based credit assignment framework that introduces counterfactual Shapley values (φ-values) into reinforcement learning for the first time. By integrating φ-values into a policy gradient algorithm—termed φ-PPO—and combining it with Prioritized Trajectory Replay (PTR), the method precisely quantifies the true causal effect of individual actions on final rewards. Evaluated under challenging conditions such as sparse causality, high environmental randomness, and delayed rewards, the approach not only achieves substantially higher attribution accuracy but also maintains optimal policy learning capability, demonstrating superior sample efficiency over existing methods and successfully solving tasks where prior state-of-the-art algorithms fail to converge.

Causal AttributionCredit Assignment ProblemEnvironmental Stochasticity

This work addresses the misalignment between short-term optimal decisions and long-term rewards in adaptive experimentation, particularly when reward shifts may occur during the commitment phase. The authors propose the RAEC algorithm, which reserves resources during the exploration phase to jointly minimize short-term regret and accurately identify the long-term optimal arm, and extend it to settings with structural priors and combinatorial commitments. Theoretically, they provide the first tight minimax characterization of the trade-off cost between short-term performance and long-term commitment, revealing that under structural priors, identifying changes in reward rankings is more critical than estimating the magnitude of shifts. They further introduce the ROSCOC algorithm, which directly maps exploration history to a committed combinatorial action. The proposed algorithms achieve tight regret upper bounds across various parameter regimes, and numerical experiments demonstrate their superiority over baseline methods.

adaptive experimentationcommitment phasepost-commitment reward shift

This work investigates Group Relative Policy Optimization (GRPO), revealing that its use of intra-group average return as a baseline imposes a zero-sum constraint on the advantage function, which undermines credit assignment and induces gradient sparsity, thereby hindering multi-step reasoning. Building upon the policy gradient theorem, the study is the first to demonstrate that the GRPO gradient matrix possesses an intrinsic rank-2 structure, proving that its effective rank remains approximately two regardless of group size. It further establishes that this baseline is optimal under specific conditions and quantifies the exacerbation of gradient sparsity during training. Through theoretical analysis, singular value decomposition, and empirical validation on Nemotron-4B with GSM8K, the paper identifies credit assignment as the critical bottleneck limiting multi-step reasoning performance.

credit assignmentgradient sparsitypolicy optimization

This work addresses the challenge in Reinforcement Learning with Verifiable Rewards (RLVR), where high-cost rollouts and excessive prompting often yield saturated groups—comprising entirely correct or incorrect responses—that provide ineffective policy gradients. To overcome this, the authors propose SARA, a method that formulates rollout collection as a sequential allocation problem under a fixed budget. SARA leverages Beta posterior estimation of prompt success rates, integrates a closed-form validity predictor, and employs a dual-threshold SPRT-style decision rule to dynamically terminate uninformative groups and reallocate resources. Theoretically, SARA guarantees reduced rollout consumption and improved returns within a fixed budget. Empirically, on mathematical reasoning and planning tasks using 1.5B/3B models on a single GPU, SARA reduces rollouts by 22% compared to dynamic sampling; when combined with DPS, it achieves slightly higher accuracy than DS while saving 67% of rollouts.

compute efficiencypolicy gradientreinforcement learning

Hot Scholars

MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
YF

Yuchen Fan

Shanghai AI Laboratory & Shanghai Jiao Tong University
NLPLarge Language ModelsEvaluation
BH

Bingxiang He

Second year PhD Candidate, Tsinghua University
Natural Language Processing
YL

Yitong Li

Huawei Technologies Co., Ltd.
Natural Language ProcessingMachine Learning