Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited predictability of latent spaces in existing world models, which stems from the decoupling of representation learning and dynamics prediction. We propose an end-to-end joint training framework that integrates vision foundation models with flow matching generative models to synergistically optimize the latent encoder and the generative dynamics model, thereby directly shaping representations amenable to temporal prediction. Furthermore, a collapse-prevention mechanism is introduced to eliminate the reliance on two-stage training pipelines. The proposed approach significantly enhances long-horizon temporal coherence, consistently outperforming existing baselines across multi-task and high-resolution scenarios.
📝 Abstract
Predicting the future evolution of a scene is a fundamental capability for world modeling. Recent work has shown that operating in the feature space of Vision Foundation Models (VFMs) yields semantically rich representations that support diverse future scene understanding tasks. However, existing approaches rely on two-stage pipelines, where VFM features are first compressed using fixed dimensionality reduction (e.g., PCA) or independently trained autoencoders, and a separate predictor is trained on top of the resulting frozen latent space. This decoupling between representation learning and temporal prediction, as well as approaches that apply predictors directly on raw VFM features, provides no guarantee that the latent space is structured for predictable dynamics. In this work, we propose Latent-Foresight, an end-to-end framework that jointly learns a latent tokenizer and a flow-based generative dynamics model, explicitly shaping the representation to support temporal predictability. To enable stable joint optimization, we introduce several key design choices that prevent latent collapse and align reconstruction with generative objectives. Extensive experiments show that our approach learns more temporally coherent latent representations and consistently outperforms two-stage baselines across multiple future scene understanding tasks and prediction horizons, while eliminating separate training stages, including during high-resolution adaptation. We provide the implementation code and model weights at https://github.com/Sta8is/Latent-Foresight
Problem

Research questions and friction points this paper is trying to address.

World Models
Latent Representations
Future Scene Prediction
Vision Foundation Models
Temporal Predictability
Innovation

Methods, ideas, or system contributions that make the work stand out.

End-to-End Learning
Latent World Models
Flow-based Generative Model
Vision Foundation Models
Temporal Predictability
🔎 Similar Papers
No similar papers found.