🤖 AI Summary
This paper addresses the problem of predicting trajectories of surrounding vehicles in autonomous driving. We propose a multimodal prediction method that jointly models static road infrastructure and dynamic traffic participants. Our approach introduces three key contributions: (1) a lane-graph-guided target-conditioned modeling mechanism that jointly encodes the target vehicle’s state and road topology; (2) a cross-context attention module that adaptively fuses heterogeneous inputs—including lane graphs, neighboring vehicle states, and historical trajectories; and (3) a hybrid architecture combining graph convolutional networks (GCNs) for lane-structure encoding, gated recurrent units (GRUs) for temporal dynamics modeling, and a Laplacian mixture density network (L-MDN) for multimodal probabilistic output. Evaluated on the nuScenes motion prediction benchmark, our method achieves state-of-the-art performance, outperforming prior work in both average displacement error (ADE) and final displacement error (FDE). Notably, it demonstrates superior robustness and physical plausibility in complex scenarios such as intersections and lane changes.
📝 Abstract
Predicting future trajectories of surrounding vehicles heavily relies on what contextual information is given to a motion prediction model. The context itself can be static (lanes, regulatory elements, etc) or dynamic (traffic participants). This paper presents a lane graph-based motion prediction model that first predicts graph-based goal proposals and later fuses them with cross attention over multiple contextual elements. We follow the famous encoder-interactor-decoder architecture where the encoder encodes scene context using lightweight Gated Recurrent Units, the interactor applies cross-context attention over encoded scene features and graph goal proposals, and the decoder regresses multimodal trajectories via Laplacian Mixture Density Network from the aggregated encodings. Using cross-attention over graph-based goal proposals gives robust trajectory estimates since the model learns to attend to future goal-relevant scene elements for the intended agent. We evaluate our work on nuScenes motion prediction dataset, achieving state-of-the-art results.