🤖 AI Summary
This study addresses the bottleneck in task specification for latent-space planning with language instructions, which arises from cross-modal noise and reliance on large models. To this end, we propose LAGO, a hierarchical world model built upon the Joint Embedding Predictive Architecture (JEPA). By employing a single regression objective, LAGO jointly predicts dynamics and grounded language within a shared latent space, decomposing long-horizon instructions into sequences of intermediate latent subgoals. It further enables flexible planning without rigid constraints through a soft minimum alignment cost. Experimental results demonstrate that LAGO significantly outperforms flat and image-goal planners in navigation and manipulation tasks. Notably, it more than doubles the long-horizon success rate compared to baselines and surpasses vision-language reward models.
📝 Abstract
Planning with world models is bottlenecked by compounding prediction errors and the difficulty of defining optimizable goals. Visual targets provide precise local gradients but poor distant guidance, while language is flexible yet limited by noisy cross-modal alignment or dependence on large generative models unsuited for the high-sampling nature of model-based planning. To address these challenges, we introduce Latent Goal Prediction from Language (LAGO), a framework that predicts both sequences of intermediate goal states from language instructions and action-conditioned rollouts, all within the same latent space. Rather than optimizing toward a single global objective, LAGO dynamically decomposes instructions into explicitly predicted, locally tractable latent subgoals. By updating these subgoals online and using a soft minimum trajectory cost during planning, LAGO enables an agent to follow coherent latent trajectories over long horizons. Evaluation across multiple environments planning horizons shows that LAGO avoids the sharp degradation of prior methods. By achieving robust and precise long-horizon planning purely from language, LAGO bridges the precision of visual goals with the flexibility of text-guided control.