Score
Designs and implements reinforcement learning training methods that use a learned discriminator's outputs (e.g., logits or scores) as the reward or shaping signal to guide a generator or policy. This includes building KL-regularized policy/objective updates and procedures that optimize the model on its own samples to improve perceptual or semantic realism.
This paper addresses the misalignment between reward models and true objectives in deep reinforcement learning, as well as the resulting limitations in policy optimization. To this end, it introduces— for the first time—a unified taxonomy that systematically organizes reward modeling across three orthogonal dimensions: modeling source (explicit vs. implicit), mechanism design (supervised vs. interactive), and learning paradigm (static vs. dynamic). The survey comprehensively covers mainstream approaches—including inverse reinforcement learning, preference learning, language-model-based feedback, human demonstration distillation, contrastive learning, and online interactive modeling—and critically analyzes evaluation methodologies and practical deployment challenges. This work fills a critical gap in the literature by providing the first systematic, cross-cutting review of reward modeling. It clarifies the technical evolution of the field and identifies four key research frontiers: scalability, generalization, robustness, and human-AI alignment.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
To address slow convergence in reinforcement learning under sparse rewards, this paper proposes a bootstrapped potential-based reward shaping (V-PBRS) method grounded in state-value function estimation. Unlike conventional approaches, V-PBRS eliminates the need for handcrafted potential functions by dynamically constructing them directly from the current estimate of the state-value function. It rigorously preserves the optimality of the original policy and provides theoretical convergence guarantees. By integrating the potential function into both Q-learning and DQN frameworks—leveraging Monte Carlo or temporal-difference online estimation mechanisms—V-PBRS achieves stable and adaptive reward shaping. Empirical evaluation on the Atari benchmark demonstrates significantly accelerated training convergence, validating its effectiveness, robustness, and generalization capability in high-dimensional visual environments.
In reinforcement learning, reward function design faces critical challenges including delayed signals, ambiguity, misalignment with task objectives, and induction of undesirable behaviors. To address these, this paper proposes three novel reward mechanisms: teacher-driven, adaptive explainable, and agent-autonomous reward generation. Our core contributions are the first-ever adaptive explainable reward design method and a meta-learning–driven autonomous reward generation framework—enabling a paradigm shift from expert-guided reward specification to online inverse reward modeling by the agent. Technically, we integrate reward shaping, eXplainable AI (XAI)-informed reward modeling, policy-value alignment, and online inverse reward design. Experiments across multiple sparse-reward benchmarks demonstrate over 40% faster training convergence, significantly improved policy robustness, and high reward interpretability—validated by domain experts with 92% inter-rater agreement.
Designing reward functions for reinforcement learning (RL) agents in games traditionally relies heavily on domain expertise and struggles to adapt to dynamic content changes. Method: This paper proposes an LLM-based automated iterative reward weight optimization method that takes user-specified behavioral objectives as input and leverages agent training feedback—such as success rate and episode length—to perform closed-loop, multi-round LLM reasoning for reward weight self-calibration, eliminating manual intervention. Contribution/Results: To our knowledge, this is the first work to integrate LLMs into online adaptive optimization of RL reward functions, substantially reducing dependence on human experts. Evaluated on a racing task, the approach improves agent success rate from 9% to 80% and reduces average lap steps to 855—performance approaching that achieved by expert manual tuning.
To address low sample efficiency and unstable convergence in reinforcement learning caused by sparse rewards, this paper proposes an adaptive reward shaping method grounded in historical success rates. The method models state-dependent success probability as a time-varying Beta distribution—explicitly capturing epistemic uncertainty for the first time in this context. It further introduces an uncertainty-driven stochastic annealing strategy that naturally balances exploration and exploitation. For scalable, model-free, nonparametric success-rate estimation in high-dimensional continuous state spaces, the approach integrates kernel density estimation (KDE) with Random Fourier Features. Experiments demonstrate substantial improvements in sample efficiency and convergence stability on extremely sparse-reward tasks, consistently outperforming state-of-the-art reward shaping and intrinsic motivation baselines across diverse benchmarks.
该研究通过解构强化学习后训练算法,探讨了其在提升大型语言模型能力时的机制和影响因素,如奖励信号、提示分布等,并分析了这些选择如何相互作用以影响后训练的成功。
论文解决了目标条件强化学习在稀疏奖励下的长时序问题,通过提出Locally-Guided Actor Critic方法,利用子目标意识的批评者来指导策略训练。
This study addresses the computational expense and instability of retraining from scratch when reward functions change during reinforcement learning (RL) post-training of foundation models. To this end, we propose a zero-shot policy prediction method that eliminates the need for additional RL training. Leveraging the low-rank subspace property of log-policies across different rewards, our approach directly estimates target policy weights under new reward functions through linear combinations of log-probabilities from existing policies. Experiments demonstrate that this method efficiently and accurately predicts RL post-training outcomes across both synthetic and real-world reward scenarios in text and image modalities. Consequently, it significantly reduces the computational overhead associated with multi-reward alignment.
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.
This study addresses the issues of policy staleness bias in asynchronous reinforcement learning for large language models and the high sampling costs associated with GRPO by proposing the KLPO framework. This method anchors the KL regularization term to the sampling policy, enabling importance-weight-free updates through least-squares fitting of log-ratios. Furthermore, it derives a closed-form Gibbs solution that eliminates the need for partition function learning and group response dependencies, demonstrating that various existing algorithms constitute special cases thereof. Notably, KLPO supports single-trajectory, critic-free updates while strictly guaranteeing unbiased gradients, thereby substantially reducing computational overhead.