WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between visual world models and raw geometric trajectory representations that constrains autonomous driving planning. To this end, we propose a method for constructing a compact trajectory latent space via a dual-branch autoencoder. This approach extracts action-relevant semantics without fine-tuning the frozen world model. By integrating JEPA, REPA, and generative trajectory latent space learning techniques, it achieves efficient alignment between visual semantics and trajectory representations. Experimental results demonstrate that our method attains a PDMS of 89.8 on the NAVSIM benchmark while reducing planner computational cost by 30.5%, thereby achieving synergistic improvements in both performance and efficiency.
📝 Abstract
Driving world models learn rich predictive representations of the surrounding environment from visual observations, yet accurate visual prediction does not necessarily translate into effective trajectory planning. We argue that a key bottleneck lies in the mismatch between visual world states and raw geometric trajectories, which may limit the planner's ability to exploit action-relevant semantics encoded by the world model. To address this issue, we propose World-Model Alignment for Latent Trajectories (WALT), which learns a compact generative trajectory latent space by transferring information from a frozen pretrained driving world model without modifying the world model itself. Rather than directly generating raw waypoints, WALT maps them into compact representations through a dual-branch trajectory autoencoder and transfers semantic knowledge from the frozen visual world model into this trajectory space, encouraging the learned action representation to capture scene-level cues relevant to future motion and planning. Beyond our proposed formulation, we systematically study latent learning based on Joint-Embedding Predictive Architectures (JEPA) and feature alignment following Representation Alignment (REPA) to investigate how trajectory-only representation learning affects downstream planning. We evaluate WALT on the NAVSIM benchmarks. Relative to the raw-waypoint baseline, WALT improves PDMS from 89.4 to 89.8 on NAVSIMv1 and EPDMS from 87.3 to 87.9 on NAVSIMv2 while reducing trajectory planner FLOPs by 30.5%. These results suggest that preserving world representations while extracting action-relevant information provides an effective interface for world-model-based trajectory planning.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Driving
World Models
Trajectory Planning
Representation Mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Trajectories
World Model Alignment
Autonomous Driving
Trajectory Autoencoder
Representation Learning
M
Mingkai Jia
The Hong Kong University of Science and Technology; Horizon Robotics
J
Jiaxin Guo
The Chinese University of Hong Kong
Z
Zhijian Shu
Horizon Robotics; Nanjing University of Posts and Telecommunications
Jiawei Xu
Jiawei Xu
Nankai University
Computer VisionMachine Learning3DVComputer Graphics
M
Mingxiao Li
Horizon Robotics
J
Jintao Cheng
The Hong Kong University of Science and Technology
Ping Tan
Ping Tan
Hong Kong University of Science and Technology (HKUST)
Computer VisionComputer Graphics
Wei Yin
Wei Yin
Staff Research Scientist, Horizon Robotics
World ModelGenerative AIPhysical AI