Flow Policies as Actions of Skill-Level World Models: Learned and Symbolic Abstractions for Long-Horizon Planning

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of large control-level action search spaces, error accumulation, and reliance on manual annotations for symbolic skills in long-horizon planning. To this end, it proposes a flow matching-based latent world model coupled with a hierarchical planning framework. By designing four skill abstraction mechanisms that map noise seeds to symbolic labels, the method automatically generates skill actions without domain knowledge, enabling single-execution state transitions that effectively integrate representation learning with symbolic reasoning. Evaluated on simulated block rearrangement tasks, the proposed unsupervised approach maintains high success rates across multi-skill scenarios, validating the effectiveness of the introduced abstraction mechanisms.
📝 Abstract
Latent world models enable robots to plan by predicting the consequences of actions. Planning long tasks with control-rate actions requires many prediction steps, which enlarges the search space and accumulates error. Skill-level actions shorten these sequences, but a symbolic skill vocabulary requires domain knowledge and labeled demonstrations. We construct skill-level actions from the inputs of a flow-matching policy trained on demonstrations segmented into complete skills. The policy maps a noise seed and an observation, optionally with a code or label, to a complete skill execution, so one execution is one world-model transition. On this mechanism we propose four action abstractions with increasing task knowledge: a compressed seed, two discrete codes learned from the demonstrations, and a symbolic label. We evaluate them with a common world-model training procedure and planning framework on simulated block rearrangement tasks that require up to 14 sequential skills. The symbolic label succeeds in over 90% of the tasks that require up to six skills and degrades beyond. Without any label, an object-centric learned code matches it on single-skill tasks and retains half to three quarters of its success on tasks of two to five skills. Ablations attribute much of the label's advantage to its planner knowing which actions are applicable, rather than to the label itself. Beyond six skills the search, not the world model, limits success. Project page: https://andreumatoses.github.io/research/flow-skill-wm
Problem

Research questions and friction points this paper is trying to address.

long-horizon planning
world models
skill-level actions
error accumulation
search space
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Flow Matching
Skill-level Abstractions
Long-Horizon Planning
Robot Manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Andreu Matoses Gimenez
Andreu Matoses Gimenez
PhD candidate at Cognitive Robotics, TU Delft
roboticstask and motion planningmulti agent systemsRL
A
Andrei-Carlo Papuc
Cognitive Robotics Department, Delft University of Technology, Delft, The Netherlands
C
Chris Pek
Cognitive Robotics Department, Delft University of Technology, Delft, The Netherlands
Javier Alonso-Mora
Javier Alonso-Mora
Associate Professor, Delft University of Technology
RoboticsIntelligent TransportationMotion PlanningMulti-Robot SystemsArtificial Intelligence