Efficient Virtuoso: A Latent Diffusion Transformer Model for Goal-Conditioned Trajectory Planning

📅 2025-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Autonomous vehicle trajectory prediction must simultaneously achieve high fidelity, diversity, and real-time performance—yet existing methods struggle to balance these three objectives. This paper proposes a goal-conditioned latent diffusion-Transformer hybrid model. First, we construct a geometry-preserving latent space via PCA dimensionality reduction and introduce a two-stage normalization strategy to ensure training stability. Second, a Transformer-based StateEncoder fuses multi-source scene context, while an MLP-based denoiser efficiently models the diffusion process in the low-dimensional latent space. Third, sparse multi-step route guidance is incorporated to explicitly encode fine-grained driving behaviors. Evaluated on the Waymo Open Motion dataset, our method achieves a minADE of 0.25—substantially outperforming state-of-the-art approaches—while maintaining computational efficiency suitable for real-time deployment.

Technology Category

Computer Vision: Diffusion Models for VisionPlanning, Routing, and Scheduling: Planning with Language ModelsHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsResponsible Web: Machine-in-the-loop, human agency and autonomyUser Modeling, Personalization and Recommendation: Accountability, Transparency, and Ethics for personalization
📝 Abstract
The ability to generate a diverse and plausible distribution of future trajectories is a critical capability for autonomous vehicle planning systems. While recent generative models have shown promise, achieving high fidelity, computational efficiency, and precise control remains a significant challenge. In this paper, we present the extbf{Efficient Virtuoso}, a conditional latent diffusion model for goal-conditioned trajectory planning. Our approach introduces a novel two-stage normalization pipeline that first scales trajectories to preserve their geometric aspect ratio and then normalizes the resulting PCA latent space to ensure a stable training target. The denoising process is performed efficiently in this low-dimensional latent space by a simple MLP denoiser, which is conditioned on a rich scene context fused by a powerful Transformer-based StateEncoder. We demonstrate that our method achieves state-of-the-art performance on the Waymo Open Motion Dataset, reaching a extbf{minADE of 0.25}. Furthermore, through a rigorous ablation study on goal representation, we provide a key insight: while a single endpoint goal can resolve strategic ambiguity, a richer, multi-step sparse route is essential for enabling the precise, high-fidelity tactical execution that mirrors nuanced human driving behavior.
Problem

Research questions and friction points this paper is trying to address.

Generating diverse autonomous vehicle trajectories with high fidelity
Achieving computational efficiency in goal-conditioned trajectory planning
Resolving strategic ambiguity through improved goal representation techniques
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent diffusion model for goal-conditioned planning
Two-stage normalization pipeline for stable training
Transformer-based StateEncoder with MLP denoiser
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Independent Researcher