🤖 AI Summary
Traditional ST-GCNs employ单一 temporal modules—either CNNs or LSTMs—leading to insufficient capture of dynamic spatiotemporal patterns. To address this, we propose a plug-and-play hybrid temporal module that, for the first time, synergistically integrates CNNs and LSTMs within a unified co-temporal block. This design jointly models local temporal features and long-range dependencies. Through theoretical analysis and cross-dataset ablation studies, we systematically characterize the intrinsic relationship between temporal module architecture and representational capacity. Evaluated on standard spatiotemporal graph benchmarks—including NTU-RGB+D and PeMSD7—our method achieves significant improvements in prediction accuracy and cross-domain generalization. It consistently outperforms pure-CNN and pure-LSTM baselines in temporal representation learning. The proposed module establishes a reusable, principled design paradigm for temporal modeling in ST-GCNs, advancing both expressiveness and architectural flexibility.
📝 Abstract
Spatio-Temporal graph convolutional networks were originally introduced with CNNs as temporal blocks for feature extraction. Since then LSTM temporal blocks have been proposed and shown to have promising results. We propose a novel architecture combining both CNN and LSTM temporal blocks and then provide an empirical comparison between our new and the pre-existing models. We provide theoretical arguments for the different temporal blocks and use a multitude of tests across different datasets to assess our hypotheses.