FLEX-WAM: Flexible Block-Causal World-Action Models for Long-Horizon Imagination and Planning

πŸ“… 2026-10-04
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high computational overhead, fixed field-of-view, and long-horizon instability of existing video action models by proposing FLEX-WAM, a flexible and efficient block-causal world action model. It unifies simulation and policy reasoning while supporting variable-length contexts and infinite autoregressive generation. Methodologically, it introduces a block-causal KV cache architecture that integrates axial attention with block-wise diffusion forcing, alongside an FD-elastic mechanism for gradient balancing to preserve action responsiveness. Flow matching training is combined with Monte Carlo Tree Search for efficient planning. Experiments demonstrate that FLEX-WAM significantly improves multi-step prediction quality and latency, achieves pure imagination-based planning in simulated benchmarks, and enables real-time error detection with self-improvement on physical robots.
πŸ“ Abstract
World--action models (WAMs) promise a unified model that predicts action-conditioned futures, generates feasible actions, and supports planning in imagination. However, existing joint video--action models often use computationally heavy, fixed-horizon backbones ill-suited to streaming inference and stable long-horizon open-loop rollouts. We introduce FLEX-WAM, a Flexible and Efficient Block-Causal World--Action Model for unified simulation and policy inference. FLEX-WAM supports variable-length contexts and non-causal prediction horizons, as well as infinite autoregressive generation frame by frame or block by block. Its block-causal, KV-cacheable architecture combines axial attention and blockwise diffusion forcing to enable efficient real-time rollout and deployment-time latency--throughput tradeoffs without retraining. Joint training can nevertheless produce plausible futures that weakly respond to commanded actions. We address this failure mode by balancing state and action flow-matching gradient contributions across the state--action diffusion-noise grid and regulating world-model sampling using Forward-Dynamics (FD) elasticity, an efficient training-time proxy for action responsiveness. Across simulated and real-world datasets, FLEX-WAM achieves superior multi-step prediction quality and latency while producing stable joint state--action rollouts for thousands of steps. As a joint action proposer and simulator within MCTS, it solves long-horizon PushT and all five OGBench Puzzle-4x4 tasks entirely in imagination. On a bimanual OpenArm-based robot, a single checkpoint jointly serves as a play policy and expected-outcome predictor, enabling real-time identification and collection of model--reality mismatches for future self-improvement.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
Long-Horizon Planning
Block-Causal Architecture
Action Responsiveness
Streaming Inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Model
Block-Causal Architecture
Diffusion Forcing
Flow Matching
Monte Carlo Tree Search
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
R
R. Khorrambakht
Center for Robotics and Embodied Intelligence (CREO), New York University
J
Joseph Amigo
Center for Robotics and Embodied Intelligence (CREO), New York University
F
FΓ©lix Lebel
Center for Robotics and Embodied Intelligence (CREO), New York University
L
Leon Seetoo
Center for Robotics and Embodied Intelligence (CREO), New York University
Jean Ponce
Jean Ponce
Ecole Normale Superieure/PSL Research University
computer visionmachine learningrobotics
Zhenzhen Li
Zhenzhen Li
Bosch Center for AI
Nonconvex optimizationApplied MathematicsDistributed Deep LearningAI compression
Ludovic Righetti
Ludovic Righetti
New York University and Artificial and Natural Intelligence Toulouse Institute
Robotics