🤖 AI Summary
Autonomous vehicle trajectory prediction must simultaneously achieve high fidelity, diversity, and real-time performance—yet existing methods struggle to balance these three objectives. This paper proposes a goal-conditioned latent diffusion-Transformer hybrid model. First, we construct a geometry-preserving latent space via PCA dimensionality reduction and introduce a two-stage normalization strategy to ensure training stability. Second, a Transformer-based StateEncoder fuses multi-source scene context, while an MLP-based denoiser efficiently models the diffusion process in the low-dimensional latent space. Third, sparse multi-step route guidance is incorporated to explicitly encode fine-grained driving behaviors. Evaluated on the Waymo Open Motion dataset, our method achieves a minADE of 0.25—substantially outperforming state-of-the-art approaches—while maintaining computational efficiency suitable for real-time deployment.
📝 Abstract
The ability to generate a diverse and plausible distribution of future trajectories is a critical capability for autonomous vehicle planning systems. While recent generative models have shown promise, achieving high fidelity, computational efficiency, and precise control remains a significant challenge. In this paper, we present the extbf{Efficient Virtuoso}, a conditional latent diffusion model for goal-conditioned trajectory planning. Our approach introduces a novel two-stage normalization pipeline that first scales trajectories to preserve their geometric aspect ratio and then normalizes the resulting PCA latent space to ensure a stable training target. The denoising process is performed efficiently in this low-dimensional latent space by a simple MLP denoiser, which is conditioned on a rich scene context fused by a powerful Transformer-based StateEncoder. We demonstrate that our method achieves state-of-the-art performance on the Waymo Open Motion Dataset, reaching a extbf{minADE of 0.25}. Furthermore, through a rigorous ablation study on goal representation, we provide a key insight: while a single endpoint goal can resolve strategic ambiguity, a richer, multi-step sparse route is essential for enabling the precise, high-fidelity tactical execution that mirrors nuanced human driving behavior.