🤖 AI Summary
Addressing the challenge of jointly modeling fine-grained (pixel-level) and coarse-grained (regional-level) spatial structures while preserving long-range temporal dependencies in traffic video forecasting, this paper proposes a tile-based spatial attention-enhanced encoder-decoder LSTM framework. Our key contributions are: (1) a tile-aware spatial attention mechanism that explicitly captures multi-scale spatial hierarchies; (2) an attention-guided LSTM cell integrating convolutional features to improve trajectory modeling fidelity; and (3) an adaptive frame-sampling strategy coupled with an end-to-end trainable encoder-decoder architecture, balancing computational efficiency and robustness. Evaluated on the Traffic4Cast 2019 dataset, our method significantly outperforms both 2D/3D CNNs and standard ConvLSTM baselines, achieving superior prediction accuracy and enhanced spatiotemporal consistency.
📝 Abstract
This extended abstract describes our solution for the Traffic4Cast Challenge 2019. The key problem we addressed is to properly model both low-level (pixel based) and high-level spatial information while still preserve the temporal relations among the frames. Our approach is inspired by the recent adoption of convolutional features into a recurrent neural networks such as LSTM to jointly capture the spatio-temporal dependency. While this approach has been proven to surpass the traditional stacked CNNs (using 2D or 3D kernels) in action recognition, we observe suboptimal performance in traffic prediction setting. Therefore, we apply a number of adaptations in the frame encoder-decoder layers and in sampling procedure to better capture the high-resolution trajectories, and to increase the training efficiency.