S3: Stable Subgoal Selection by Constraining Uncertainty of Coarse Dynamics in Hierarchical Reinforcement Learning

📅 2026-07-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the instability in high-level subgoal selection within hierarchical reinforcement learning, which arises due to sparse and delayed environmental feedback and is exacerbated by the limitations of low-level execution capabilities. To mitigate this issue, the authors propose an intrinsic motivation mechanism based on coarse-grained dynamic modeling: a coarse dynamics model is constructed by aggregating multi-step environmental transitions, and a Mixture Density Network (MDN) is employed to quantify the predictive uncertainty of this model. This uncertainty is then used as a risk-sensitive intrinsic reward to guide the high-level agent away from subgoals associated with high uncertainty. Evaluated on non-stationary, long-horizon tasks, the proposed method significantly outperforms existing hierarchical reinforcement learning algorithms, demonstrating improved task completion efficiency and policy stability.
📝 Abstract
Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution. It has been widely successful in solving long-horizon and complex tasks, where flat-RL algorithms have difficulty in learning. However, while the low-level agent in HRL benefits from dense feedback and abundant trial opportunities, the high-level agent receives sparse, delayed feedback from the environment and its performance depends on the low-level execution capability. In this paper, we study whether subgoal selection by the high-level agent can be performed more strategically, by providing it with dynamics-aware intrinsic motivation. Since motivation based on primitive transition dynamics would require broad coverage of the state-action space, we propose to use coarse dynamics, i.e., environment transitions aggregated over multiple steps at the temporal scale at which the high-level agent operates. This approach stabilizes the high-level policy by learning to minimize the predictive uncertainty associated with the coarse dynamics, and provides a guided structure for navigation. We model the predictive uncertainty by evaluating different dispersion metrics as approximated by a Mixture Density Network (MDN). Empirically, we observe that a dense, dynamics-aware intrinsic reward leads to risk-averse subgoal selection, enabling it to outperform state-of-the-art HRL methods in non-stationary long-horizon environments.
Problem

Research questions and friction points this paper is trying to address.

Hierarchical Reinforcement Learning
subgoal selection
coarse dynamics
predictive uncertainty
intrinsic motivation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Reinforcement Learning
Coarse Dynamics
Predictive Uncertainty
Intrinsic Motivation
Mixture Density Network
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kshitij Kumar Srivastava
University of Massachusetts, Lowell
K
Kshitij Jerath
University of Massachusetts, Lowell