🤖 AI Summary
Addressing the challenge of simultaneously ensuring diversity, plausibility, and interpretability in vision-driven multi-trajectory prediction (MTP), this paper presents a systematic survey of the field. We propose the first unified taxonomy for MTP, comprehensively categorizing model paradigms—including variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, graph neural networks (GNNs), and attention mechanisms—alongside mainstream open-source datasets and evaluation metrics. By analyzing the interplay between uncertainty modeling and interpretability, we identify key technical bottlenecks and distill five emerging research directions. This work establishes a principled theoretical framework and practical guidelines for developing safe, reliable autonomous navigation systems. (128 words)
📝 Abstract
Trajectory prediction is an important task to support safe and intelligent behaviours in autonomous systems. Many advanced approaches have been proposed over the years with improved spatial and temporal feature extraction. However, human behaviour is naturally multimodal and uncertain: given the past trajectory and surrounding environment information, an agent can have multiple plausible trajectories in the future. To tackle this problem, an essential task named multimodal trajectory prediction (MTP) has recently been studied, which aims to generate a diverse, acceptable and explainable distribution of future predictions for each agent. In this paper, we present the first survey for MTP with our unique taxonomies and comprehensive analysis of frameworks, datasets and evaluation metrics. In addition, we discuss multiple future directions that can help researchers develop novel multimodal trajectory prediction systems.