š¤ AI Summary
To address the challenges of model complexity, training inefficiency, and high memory overhead in spatiotemporal forecasting (e.g., traffic flow, meteorological modeling), this paper proposes a lightweight spectral-decoupled knowledge distillation framework. Our method introduces a novel frequency-alignment strategy that decomposes the teacher modelās multi-scale spatiotemporal representations into low- and high-frequency components, which are then selectively transferred to a structurally simplified student network. Specifically, low-frequency componentsācapturing global trendsāare modeled by a convolutional-Transformer hybrid encoderālatent evolutionādecoder architecture, while high-frequency componentsāencoding local dynamicsāare captured via spectral analysis. Evaluated on the NavierāStokes dataset, our approach reduces MSE and MAE by 81.3% and 52.3%, respectively, achieving substantial computational efficiency gains without sacrificing prediction accuracy. This work establishes an optimal trade-off between model fidelity and inference efficiency in spatiotemporal forecasting.
š Abstract
Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation (termed SDKD), which transfers the multi-scale spatiotemporal representations from a complex teacher model to a more efficient lightweight student network. The teacher model follows an encoder-latent evolution-decoder architecture, where its latent evolution module decouples high-frequency details and low-frequency trends using convolution and Transformer (global low-frequency modeler). However, the multi-layer convolution and deconvolution structures result in slow training and high memory usage. To address these issues, we propose a frequency-aligned knowledge distillation strategy, which extracts multi-scale spectral features from the teacher's latent space, including both high and low frequency components, to guide the lightweight student model in capturing both local fine-grained variations and global evolution patterns. Experimental results show that SDKD significantly improves performance, achieving reductions of up to 81.3% in MSE and in MAE 52.3% on the Navier-Stokes equation dataset. The framework effectively captures both high-frequency variations and long-term trends while reducing computational complexity. Our codes are available at https://github.com/itsnotacie/SDKD