HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of infeasible or redundant paths caused by fixed planning horizons in offline goal-conditioned reinforcement learning. To overcome this limitation, we propose HorizonFlow, a hierarchical planner that departs from the conventional "determine-then-generate" paradigm by treating the planning horizon as a generative output rather than a predefined input. Specifically, HorizonFlow jointly optimizes sequence content and length through an insertion-based generation mechanism coupled with flow matching techniques. Partially generated plans guide token insertion, while length information is reused to preferentially select shorter paths without requiring an additional value model. Furthermore, latent-space subgoal routing enables hierarchical action control. Extensive experiments on the Maze2D, Multi2D, and OGBench benchmarks demonstrate that HorizonFlow achieves state-of-the-art average performance.
📝 Abstract
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.
Problem

Research questions and friction points this paper is trying to address.

offline goal-conditioned reinforcement learning
trajectory planning
variable-length horizon
generative planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Variable-Length Planning
Flow Matching
Insertion-based Generation
Hierarchical Planner
Offline Goal-Conditioned RL
🔎 Similar Papers
No similar papers found.