Score
Designs and analyzes algorithms, reward schedules, and training procedures that define, combine, and optimize feedback signals across multiple hierarchical levels (e.g., coarse high-level objectives and finer low-level shaping) to improve learning and credit assignment. Includes methods that progressively densify sparse, high-level rewards into multiscale guidance and perform hierarchical policy search or reinforcement learning over abstract subpolicies, temporal abstractions, or controller layers.
Sparse, delayed, and weakly informative reward signals severely hinder the efficiency of reinforcement learning, and existing reward shaping methods often fail to adapt to dynamic environments. This work proposes the first unified analytical framework encompassing temporal, informational, and theoretical dimensions to systematically categorize and compare twelve classes of dynamic reward shaping and related adaptive mechanisms. It clearly distinguishes between parameter corrections and state-dependent modifications, and precisely delineates the boundaries among additive shaping, reward replacement, and correlated guidance. Through integrated theoretical analysis and taxonomic synthesis, the study identifies conditions under which optimality is preserved in the presence of modern RL components such as experience replay, bootstrapped critics, and reward normalization, and for the first time elucidates the intrinsic relationship between adaptation rate and learning stability.
Designing reward functions for reinforcement learning in unstructured environments remains challenging due to sparse and ambiguous task specifications. Method: This paper proposes a reward modeling approach based on multi-level, episode-wise human scoring feedback—departing from conventional binary preference comparisons. It introduces a global, non-Markovian episodic scoring mechanism and formulates a unified Bayesian inference and inverse reinforcement learning framework for end-to-end co-learning of reward functions and policies. Contribution/Results: To our knowledge, this is the first work to systematically integrate multi-level episodic feedback into reward modeling. We provide theoretical guarantees, proving a sublinear regret bound for the proposed algorithm. Empirical evaluation across diverse simulated robotic and navigation tasks demonstrates significant improvements in sample efficiency and policy performance, validating the efficacy of high-information-density scoring feedback over coarse-grained alternatives.
Addressing practical challenges in reinforcement learning—such as sparse and delayed rewards and training instability—this paper presents a systematic survey of reward engineering and reward shaping. We propose the first fine-grained taxonomy of reward design techniques, explicitly exposing their implicit assumptions and failure boundaries. Furthermore, we introduce an evaluation framework for reward shaping that jointly balances interpretability and empirical effectiveness. Our analysis integrates theoretical foundations of RL, deep RL practice, formal modeling of reward functions, and cross-domain applications—including robotics and autonomous driving. This work fills a critical gap by providing the first comprehensive, methodology-driven survey of reward design. It establishes a unified tripartite research framework comprising methodology, taxonomic classification, and application boundaries. The resulting synthesis delivers a reproducible, transferable engineering guide for algorithm designers, significantly enhancing the robustness and real-world deployability of RL systems. (149 words)
In hierarchical reinforcement learning, the inseparability of execution and policy suboptimality severely hinders practical deployment. This study addresses this challenge by decoupling the two for the first time, establishing conditions for Markovian execution optimality and reformulating execution design as an explicit optimization component. Furthermore, this work proposes a unified value function based on task-execution trees alongside a four-stage generalized hierarchical Bellman equation. Experimental results demonstrate that the proposed method achieves both execution improvement and multi-stage policy enhancement at arbitrary hierarchy depths. Additionally, it validates the complementary gains of execution optimization and goal learning in stochastic environments. Overall, this research provides a systematic theoretical framework for hierarchical decision-making.
In reinforcement learning, Linear Temporal Logic (LTL) tasks suffer from sparse rewards, hindering subgoal guidance, slowing policy convergence, and degrading robustness. To address this, we propose a progress-aware adaptive reward shaping method: for the first time, we quantify LTL satisfaction progress as a continuous, differentiable reward signal and design an online mechanism to dynamically update the reward function according to the agent’s current learning state. Our approach integrates LTL task compilation, formal progress modeling, and deep RL frameworks (PPO/SAC). Evaluated across multiple benchmark environments, the method significantly accelerates convergence, increases average expected return by 23%, improves task completion rate by 31%, and outperforms both traditional sparse-reward baselines and handcrafted reward-shaping approaches in terms of robustness and generalization.
This work addresses the challenges of exploration and inefficient policy learning in sparse-reward, long-horizon tasks by proposing a two-level hierarchical reinforcement learning framework. The high-level controller performs strategic planning to guide long-term exploration, while the low-level policy leverages Soft Actor-Critic (SAC) for continuous control, augmented with entropy regularization to enhance both policy diversity and stability. By effectively integrating hierarchical structure with maximum-entropy learning, the proposed method significantly outperforms standard SAC baselines on the SAR-2 dataset, achieving notable improvements in task success rate, environmental coverage efficiency, and convergence speed.
This work addresses the poor generalization of reinforcement learning (RL) when training and testing environments exhibit distributional shifts—a limitation exacerbated in privacy-sensitive or data-scarce settings where diverse training environments and full trajectory access are unavailable. To overcome this, the authors propose GERS, a novel method that, for the first time, integrates evolutionary algorithms with RL under the constraint of accessing only scalar feedback from validation environments without observing their trajectories. GERS employs a bilevel optimization framework: the lower level learns a policy from limited training environments, while the upper level leverages CMA-ES to optimize reward-shaping parameters for enhanced generalization. Experiments demonstrate that GERS significantly outperforms standard RL baselines across multiple continuous control tasks and achieves generalization performance comparable to domain randomization methods—despite not requiring access to trajectory data.
Existing hierarchical decision-making approaches often struggle to simultaneously satisfy constraints and maintain computational efficiency due to misalignment between low-level policies and high-level objectives. This work proposes a principled inverse optimization–based hierarchical framework that, for the first time, systematically constructs structured low-level optimization problems from expert demonstrations, thereby aligning high-level task abstractions with low-level decision-making. By integrating inverse optimization, hierarchical reinforcement learning, and optimal control, the method achieves both interpretability and computational efficiency. Empirical evaluations on resource allocation and obstacle avoidance tasks demonstrate that the approach significantly outperforms end-to-end reinforcement learning, learning-augmented optimal control, and existing hierarchical methods, achieving state-of-the-art performance in both decision quality and computational speed.
论文解决了目标条件强化学习在稀疏奖励下的长时序问题,通过提出Locally-Guided Actor Critic方法,利用子目标意识的批评者来指导策略训练。
This work addresses the instability in high-level subgoal selection within hierarchical reinforcement learning, which arises due to sparse and delayed environmental feedback and is exacerbated by the limitations of low-level execution capabilities. To mitigate this issue, the authors propose an intrinsic motivation mechanism based on coarse-grained dynamic modeling: a coarse dynamics model is constructed by aggregating multi-step environmental transitions, and a Mixture Density Network (MDN) is employed to quantify the predictive uncertainty of this model. This uncertainty is then used as a risk-sensitive intrinsic reward to guide the high-level agent away from subgoals associated with high uncertainty. Evaluated on non-stationary, long-horizon tasks, the proposed method significantly outperforms existing hierarchical reinforcement learning algorithms, demonstrating improved task completion efficiency and policy stability.