🤖 AI Summary
This study addresses the limitation that latent representations in existing world models are not ordered by predictive importance, thereby constraining planning efficiency. To overcome this, we propose ALeWM, a framework built upon the JEPA architecture that learns sequence-conditioned prefix length distributions. It introduces the first adaptive latent capacity mechanism alongside MixSIGReg regularization, which guides high-value predictive information to concentrate within compact leading latent variables. We provide theoretical guarantees and practical realization of front-loading critical information, transcending the constraints of fixed-width representations. Evaluated on controlled dynamical systems and visual control tasks, ALeWM achieves superior success rates compared to fixed-width baselines while requiring a lower average planning capacity.
📝 Abstract
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.