Score
Design and implement planners that reason over belief states using hierarchical abstractions and macro-actions to produce multi-step sequential plans. Build rollout and evaluation mechanisms (e.g., finite-horizon rollouts on hierarchical structures such as an HSG) that fuse prior knowledge with online evidence to estimate long-term expected returns and select globally consistent actions that reduce backtracking.
This study addresses the definition, discovery mechanisms, and applicability boundaries of “high-quality temporal structure” in hierarchical reinforcement learning (HRL). To tackle long-horizon dependencies, high environmental dynamics, and compositional task structures in complex open-world settings, we propose the first unified analytical framework for HRL benefits grounded in the intrinsic computational hardness of sequential decision-making. We introduce a taxonomy of temporal abstraction that spans online/offline learning and LLM-augmented paradigms, integrating hierarchical RL, LLM-guided policy decomposition, and option discovery. Our core contributions are threefold: (i) a formal definition of high-quality temporal structure—characterized by improved exploration efficiency, generalization, and interpretability; (ii) an identification of its optimal applicability in highly dynamic, long-horizon, and compositional task domains; and (iii) a systematic characterization of the fundamental trade-offs among exploration, generalization, and interpretability in HRL.
This work addresses the challenges of large action spaces, high stochasticity, long decision horizons, and resource constraints in sequential stochastic combinatorial optimization by proposing a model-based hierarchical reinforcement learning framework. The approach integrates a world model formulated as a semi-Markov decision process (SMDP) with a latent-space tree-search planner, constructing temporal structures of abstract actions through multi-timescale objectives and jointly learning budget-aware policies conditioned on subgoals to enable context-sensitive resource allocation. A key innovation lies in incorporating adaptive temporal abstraction into hierarchical planning, allowing the latent dynamics to explicitly capture the effective duration of abstract actions. Experimental results demonstrate that the proposed method significantly outperforms strong existing baselines across multiple challenging benchmarks.
In multi-objective sequential decision-making, conventional goal-conditioned (GC) policies—optimized solely for the current goal—often cause path blockages and render subsequent goals unreachable due to neglect of inter-goal dependencies. Method: We propose Dual-Objective Conditional MDPs and Two-Step Lookahead Goal-Conditioned MDPs, the first frameworks to explicitly model sequential goal dependencies within the GC policy’s conditioning mechanism. Built upon the TD3+HER architecture, our approach establishes a novel hierarchical RL paradigm that jointly optimizes for both the current and subsequent goal reachability during policy training. Contribution/Results: Evaluated on navigation and inverted pendulum tasks, our method significantly improves policy stability and sample efficiency, consistently outperforming standard GC-MDPs and single-objective GC baselines. It effectively overcomes the local optimization limitation inherent in traditional GC policies by incorporating lookahead goal dependencies into the conditioning structure.
To address the limitations of conventional distance-based methods in long-horizon visual planning—specifically their inability to model long-range dependencies and enable efficient re-planning within hierarchical reinforcement learning (HRL)—this paper proposes Discrete Hierarchical Planning (DHP). DHP achieves end-to-end hierarchical decision optimization via recursive generation of discrete abstract subgoals, tree-structured trajectory advantage estimation (which implicitly favors short-horizon plans while enabling ultra-deep generalization), and on-policy imagined-data-driven SAC training with active exploration. Key innovations include: (i) the first discrete subgoal planning paradigm for visual HRL; (ii) a tree-aware advantage estimator; and (iii) a closed-loop imagination–exploration co-training mechanism. Evaluated on a 25-room long-horizon visual navigation task, DHP significantly improves success rate, reduces average episode length, and lowers planning complexity to O(log N). Ablation studies confirm the critical contribution of each component.
This work addresses the limitations of existing large language model (LLM) agents, which typically employ fixed-granularity planning mechanisms that struggle to balance efficiency on simple tasks with the detailed reasoning required for complex ones. To overcome this, the paper introduces AdaPlan-H, a cognitively inspired adaptive hierarchical planning framework that, for the first time, integrates a progressive refinement strategy into LLM agents to enable dynamic, task-difficulty-aware adjustment of planning granularity. By synergistically combining hierarchical task decomposition, imitation learning, and capability enhancement, AdaPlan-H supports the adaptive generation and continuous optimization of planning hierarchies. Experimental results demonstrate that AdaPlan-H significantly improves success rates on multi-step complex tasks while effectively avoiding over-planning, thereby validating its efficiency and flexibility.
To address challenges in complex planning tasks—including excessively long reasoning chains, diverse constraints, and heterogeneous subtasks—this paper proposes the hypertree-structured planning paradigm. Methodologically, it adopts a divide-and-conquer strategy to enable hierarchical and dynamic reasoning: (i) it introduces the hypertree structure to model multi-granularity planning processes with constraint awareness and scalable inference; (ii) it designs an autonomous iterative refinement framework that overcomes expressivity limitations of conventional linear or tree-structured planners; and (iii) it develops an LLM-based dynamic hypertree expansion mechanism, a constraint-driven subtask decomposition and coordinated scheduling algorithm, and an iterative planning outline optimization technique. Evaluated on the TravelPlanner benchmark, our approach achieves state-of-the-art accuracy using Gemini-1.5-Pro, outperforming o1-preview by 3.6× in planning performance.
Offline goal-conditioned reinforcement learning often suffers from low sample efficiency due to redundancy in state-goal pairs. This work proposes a hierarchical policy that abstracts goal conditioning away from an absolute reference frame by introducing relativized options and multi-level state representations, enabling experience reuse across similar contexts. Unlike conventional hierarchical approaches that primarily provide temporal abstraction, the proposed architecture explicitly exploits hierarchy to eliminate redundancy in the state-goal space. Experimental results demonstrate that the method substantially improves performance on offline goal-conditioned tasks, validating the effectiveness of the introduced abstract inductive bias.
Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulation-based planning is a sampling model, which is repeatedly called to generate trajectories and estimate the utilities of available actions. Among the actions thus explored, one with the maximum estimated utility is then executed. In this paper, we examine the effect of using common random numbers in the simulation process. We obtain a simple recipe for (provably) reducing variance in relative utility when simulations invoke a rollout policy beyond some depth. Experiments on synthetic tasks confirm that our scheme improves task performance. The broader significance of our innovation is apparent from two practical applications: (1) single-step lookahead planning in a pension-disbursement task, and (2) a deployment of the well-known UCT algorithm for the game of Ludo.
This work addresses the critical yet poorly understood role of long-horizon multi-turn planning in foundation model agents, which is hindered by the uncontrolled nature of internet-scale pretraining data. The authors construct a unified and controllable multi-turn environment to systematically investigate how the format, distribution, and quality of pretraining data influence planning capabilities. During post-training, they introduce GRPO, Online Policy Distillation (OPD), and Multi-teacher Online Policy Distillation (MOPD) to shape and integrate planning skills. Their findings reveal the essential role of explicit world models in enabling long-horizon generalization, demonstrate OPD’s superiority over GRPO under low-quality long-horizon data, and present the first successful fusion and transfer of cross-environment planning abilities via MOPD. Experiments further validate the efficacy of limited high-quality data and elucidate MOPD’s robust generalization, continual learning, and interference resilience across compatible, partially shared, or conflicting planning paradigms.
This work addresses the limited policy expressivity and insufficient exploration efficiency in hierarchical reinforcement learning by proposing a novel paradigm in which a controller constructs a flexible behavioral space through linear combinations of multiple option reward functions. Departing from conventional assumptions of long-horizon planning, the approach demonstrates that the core advantage of hierarchical structures lies in their capacity to enhance exploration. Empirical evaluation in the NetHack Learning Environment shows that the proposed method substantially improves both exploration efficiency and overall learning performance, thereby validating its effectiveness and superiority in complex environments.