π€ AI Summary
This study addresses the limited long-horizon planning capabilities of existing methods caused by fixed action chunks and short-range supervision. We propose a JEPA-based world model that jointly trains a causal encoder and an autoregressive executor using mixed-span objective supervision and variable-length action chunks, incorporating Student Forcing to mitigate exposure bias. Furthermore, we introduce the ARCEM algorithm, which combines residual search with latent-space prediction to enable training-free dynamic adjustment of planning chunk lengths. Evaluated across four benchmarks, our approach achieves an average success rate of 89.29%, surpassing the strongest baselines. Additionally, long-chunk planning accelerates inference speed by approximately 1.3Γ while maintaining high success rates.
π Abstract
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.