Beyond a single latent space: a dual-latent world model for long-horizon planning

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited goal discriminability in latent world models during long-horizon planning, caused by recursive error accumulation and high-dimensional distance concentration. To overcome this, we propose a dual-latent hierarchical world model that decouples local execution from global planning, enabling reliable long-horizon control through macro-action learning. Furthermore, we introduce a dual-timescale weighted rolling supervision mechanism that optimizes multi-step prediction consistency via exponentially decaying weights, substantially enhancing the informativeness of goal representations. Integrated with end-to-end visual reinforcement learning, the proposed framework achieves average success rates of 84.4% and 69.5% under 50-step and 100-step goal offsets, respectively, across five visual control tasks, significantly outperforming existing baselines.
📝 Abstract
Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation Learning with Weighted Rollout (LoRe), which supervises self-generated predictions at both levels. An analysis of recursive error propagation motivates exponential horizon weights with separate decay rates for the two temporal scales. During planning, the high-level model generates latent subgoals that the low-level model refines into actions for precise execution. We evaluate from-scratch Dual-WM on five goal-conditioned visual control tasks against the task-wise strongest baselines without actor-guided proposals. At goal offsets of 50 and 100 environment steps, mean success increases from 75.9% to 84.4% and from 61.4% to 69.5%, respectively. At offset 100, Dual-WM outperforms these baselines on all five tasks and improves mean success over LeWM by 30.8 percentage points. Ablations and supporting analyses provide evidence of more informative representations for goal evaluation and greater consistency under recursive prediction. These results highlight the value of separating temporal roles and training across multiple horizons for reliable latent planning. Our core implementation is available at https://github.com/DeLin1001/Dual-WM-Official.
Problem

Research questions and friction points this paper is trying to address.

latent world model
long-horizon planning
error accumulation
distance concentration
goal discrimination
Innovation

Methods, ideas, or system contributions that make the work stand out.

dual-latent world model
long-horizon planning
macro-actions
weighted rollout
representation learning