🤖 AI Summary
This study addresses the issue of noise redundancy interfering with action prediction in reconstruction-based world models by proposing a decoder-free bidirectional Transformer architecture built upon a JEPA latent space. Methodologically, efficient state alignment is achieved through end-to-end joint training across four dynamic modes. The decoder-free design eliminates visual redundancy, while model predictive control (MPC) planning is executed by integrating diffusion guidance with flow matching within the policy head's noise space, effectively mitigating dynamics errors. Experimental results demonstrate that the proposed approach significantly improves state reading accuracy and planning robustness, achieving closed-loop control performance comparable to that of dedicated policy models.
📝 Abstract
World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.