WorldGuide: Goal-Directed Video World Model for Procedural Task Execution

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of closed-loop planning-execution coupling in video generation models for long-horizon procedural tasks by proposing a closed-loop visual world model. It pioneers formulating procedural video generation as closed-loop execution within the visual space: predicting atomic actions and generating videos solely from an initial image and goal, while leveraging generation feedback to drive adaptive decision-making. A hierarchical visual memory mechanism maintains long-horizon state consistency, and a planner-executor architecture is jointly trained on shared demonstration data to achieve end-to-end goal-directed control. Additionally, WorldGuide-Bench, a dataset with 59K step-level annotations, is constructed. Experiments demonstrate that the proposed method achieves a 33.33% success rate on WorldGuide-Bench, outperforming MiniMax-H3, and reaches 47.69% under goal-only conditions on VideoCraft-Bench, significantly surpassing existing approaches.
πŸ“ Abstract
Video generators and video-based world models can synthesize plausible visual trajectories, but long-horizon procedural tasks require generation to adapt to what has actually been produced. A model must determine the next action from its generated state, execute that action, and recognize when the task is complete. Open-loop generation cannot adapt to execution outcomes, while existing closed-loop systems often rely on pretrained executors or indirect verification. This leaves a gap between deciding an action and successfully realizing it. We formulate procedural video generation as \emph{closed-loop task execution in visual world space} and introduce \textbf{WorldGuide}. Given only an initial image and a task goal, WorldGuide predicts an atomic action, generates its corresponding video clip, and uses the generated result to select the next action or terminate. The Planner and Executor are trained on the same step-level procedural demonstrations: the Planner learns to predict the next atomic action or task completion from visual progress, while the Executor is directly trained to realize the predicted actions. Hierarchical visual memory maintains state across long-horizon execution with bounded history token cost. Due to the lack of step-level action-video supervision for joint planner-executor training, we introduce \textbf{WorldGuide Bench}: approximately 59K step-annotated videos across 245 tasks and 27 procedural categories. WorldGuide achieves a 33.33\% Task Success on \textbf{WorldGuide-Bench}, compared with 29.90\% for the strong recent video model MiniMax-H3, even though MiniMax-H3 receives reference action plans, and achieves 47.69\% on \textbf{VideoCraft-Bench} compared with 32.73\% for MiniMax-H3 under goal-only conditioning. These results demonstrate the importance of coupling planning with learned execution for goal-directed procedural video generation.
Problem

Research questions and friction points this paper is trying to address.

Video World Model
Procedural Task Execution
Closed-loop Generation
Goal-directed Planning
Long-horizon Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Closed-loop Video Generation
World Model
Procedural Task Execution
Hierarchical Visual Memory
Joint Planner-Executor Training
πŸ”Ž Similar Papers
No similar papers found.