Score
Designs and implements mechanisms that rescale or reweight reward and advantage signals during training—using adaptive importance weighting, dynamic reward rescaling, discriminative or group-variance weighting—to stabilize policy-gradient or likelihood-based updates and reduce variance in credit assignment. Builds online estimators and update rules that adjust reward magnitudes or per-group/per-sample weights to balance learning emphasis, mitigate noisy gradients, and improve exploration and optimization stability.
This work addresses the pervasive issue of length inflation in large language models trained with reinforcement learning, where reward-driven optimization often leads to excessively verbose outputs without compromising downstream performance. To tackle this challenge, the authors propose Group Relative Reward Rescaling (GR³), a novel framework that introduces the first general, continuous, and reward-dependent length gating mechanism. GR³ dynamically adapts to instance difficulty and preserves high-quality trajectory signals through multiplicative reward rescaling, group-relative regularization, and advantage-aware calibration. Seamlessly integrated into the GRPO training pipeline, the method significantly mitigates length inflation under both RLHF and RLVR paradigms, outperforming current state-of-the-art baselines while maintaining or even enhancing downstream task performance.
In reinforcement learning, adaptive interaction data—where the behavior policy is nonstationary—invalidates standard estimators, undermining asymptotic normality for off-policy counterfactual policy evaluation and dynamic treatment effect (DTE) inference. To address this, we propose a weighted Z-estimation framework that constructs time-varying adaptive weights to stabilize heteroskedasticity, achieving, for the first time in the RL off-policy setting, both consistent and asymptotically normal DTE estimation. Our approach integrates dynamic causal inference with asymptotic statistical theory, enabling rigorous hypothesis testing and construction of uniformly valid confidence regions. Simulation studies and real-world RL experiments demonstrate substantial improvements in confidence interval coverage and statistical power. The method provides the first solution for structural parameter inference under adaptive experimentation that simultaneously offers theoretical guarantees—namely consistency, asymptotic normality, and uniform validity—and empirical robustness.
Reinforcement learning agents often exhibit poor generalization to unseen environments due to overfitting to training conditions. To address this, we propose the first method that predicts agent generalization performance directly from neural network weight signals and integrates this prediction into the PPO objective for generalization-aware policy optimization. Our key contributions are: (1) a differentiable weight-feature extraction module that maps model parameters to a scalar generalization score; and (2) a generalization-aware regularization term incorporated into the PPO loss, which explicitly encourages learning of robust, environment-invariant representations. Experiments across diverse generalization benchmarks—including visual-observation domains (ProcGen) and dynamics-shift settings (MultiRoom)—demonstrate substantial improvements in cross-environment performance: our method achieves an average generalization score 23.6% higher than standard PPO, without requiring environmental augmentation, domain randomization, or auxiliary supervision.
Existing scalarization approaches in multi-reward reinforcement learning, such as reward or advantage aggregation, often suffer from training instability due to neglecting inter-objective correlations. This work proposes Dynamic Variance-Adaptive Advantage Optimization, which operates within the Group Relative Policy Optimization framework and dynamically adjusts combination weights based on the empirical variance of each objective within rollout groups. By amplifying signals from high-confidence objectives and suppressing noisy ones, and further incorporating adaptive cross-objective regularization to bound advantage magnitudes, the method ensures stable optimization without requiring a value model. Evaluated on mathematical reasoning and tool-use tasks with Qwen3 and Qwen2.5, it significantly outperforms baseline methods, achieving superior Pareto fronts and enhanced training stability.
Autoregressive models typically require retraining to align with new reward functions when they change, leading to inefficiency. This work proposes Reward-Conditioned Classifier-Free Guidance (RCFG), which formalizes reward-guided generation as a policy improvement operator for the first time, enabling flexible optimization of arbitrary reward functions at test time without retraining. By integrating autoregressive modeling, policy distillation, and reinforcement learning, RCFG supports zero-shot reward adaptation and serves as an effective warm-start to accelerate subsequent reinforcement learning convergence. Evaluated on molecular design tasks, RCFG demonstrates substantial improvements in both test-time reward optimization capability and training efficiency.
This work addresses the instability and collapse of output diversity commonly encountered in post-training with reinforcement learning, as well as the lack of a unified design principle in existing advantage function methods. The authors propose FADE, a novel framework that systematically decouples the gradient weighting structure of the advantage function by decomposing it along the sign and difficulty axes into positive and negative gradient quality components. This decomposition reveals the dynamic trade-off between exploration and exploitation, enabling an adaptive scheduling mechanism that dynamically adjusts gradient weights to balance accuracy and diversity. Evaluated on 7B and 32B models, FADE achieves peak pass@1 performance 20k and 2k training steps earlier, respectively, and demonstrates state-of-the-art accuracy–diversity trade-offs on the LiveCodeBench and AIME benchmarks.
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.
This work addresses the excessive sensitivity of existing neural reward models to semantically equivalent responses, a flaw that often triggers reward gaming in reinforcement learning and degrades policy performance. To mitigate this issue, the authors propose a training-free discretization method that leverages Monte Carlo Dropout to generate reward clusters, mapping continuous rewards to discrete values while preserving discriminative capacity. To better evaluate reward models, they introduce two novel metrics: discriminability and specificity. Empirical results demonstrate that the proposed approach significantly suppresses reward gaming and enhances policy quality across both control and natural language reinforcement learning environments.