🤖 AI Summary
This study addresses the inefficiency of planning and the issue of representation collapse in world models under high-dimensional observations. We propose a structured bilinear dynamics parameterization method based on the Joint Embedding Predictive Architecture (JEPA). By enforcing action-recoverability constraints and applying nonlinear system transformations, representation collapse is eliminated at the architectural level. Furthermore, an encoder is employed to learn structurally rich representations for optimizing control policies. This approach reduces planning time by nearly three orders of magnitude while significantly enhancing both long-horizon and real-time control accuracy, offering an efficient solution for online decision-making in complex, high-dimensional systems.
📝 Abstract
World models jointly learn latent representations and dynamics that predict how high-dimensional observations evolve under actions. In this work, we propose a JEPA-style world model in which, rather than learning arbitrary latent dynamics, we restrict them to follow a bilinear parameterization. This structure enables efficient planning and control while shifting the modeling burden onto the encoder, encouraging richer representations that expose the controllable geometry of the system. In particular, this structured parameterization allows us to structurally enforce action recoverability, thereby preventing representation collapse by construction. Although prescribing a bilinear parametrization may appear restrictive, we show that a broad class of nonlinear dynamical systems admits a transformation under which the dynamics become bilinear. Empirically, we show across standard 2D and 3D control tasks that representations with bilinear-parameterized dynamics can be learned directly from high-dimensional observations, reducing planning time by nearly three orders of magnitude while retaining or even improving control accuracy. We also propose more demanding regimes of longer-horizon planning and real-time control, and demonstrate that our method succeeds in both, moving JEPA-style world models beyond short-horizon offline planning.