🤖 AI Summary
In hierarchical reinforcement learning, the inseparability of execution and policy suboptimality severely hinders practical deployment. This study addresses this challenge by decoupling the two for the first time, establishing conditions for Markovian execution optimality and reformulating execution design as an explicit optimization component. Furthermore, this work proposes a unified value function based on task-execution trees alongside a four-stage generalized hierarchical Bellman equation. Experimental results demonstrate that the proposed method achieves both execution improvement and multi-stage policy enhancement at arbitrary hierarchy depths. Additionally, it validates the complementary gains of execution optimization and goal learning in stochastic environments. Overall, this research provides a systematic theoretical framework for hierarchical decision-making.
📝 Abstract
Hierarchical reinforcement learning uses temporally extended subtasks for exploration, yet committing to their execution can restrict both deployment and policy learning. We identify and separate the resulting execution and policy suboptimality. Task and execution trees distinguish reward objectives from policy choices and decision interruption. A Unified Value Function for HRL and a four-stage Generalized Hierarchical Bellman Equation then support a common analysis of both losses. Under bounded rewards and uniform termination, we establish hierarchical policy and execution improvement results. With the remaining node policies fixed, task-subtree compatibility and node-policy optimality under the original execution mode establish when Markov execution is optimal. The resulting decomposition leads to independent execution choices for behavior, targets, and deployment. We instantiate this principle through execution improvement and one-stage or two-stage policy improvement at arbitrary hierarchy depth. Option-based and goal-conditioned experiments demonstrate complementary gains from changing execution and changing the learning target. Controlled stochastic environments show how these gains depend on stochastic transition strength and spatial structure. This framework makes execution design an explicit component of hierarchical policy optimization.