Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

šŸ“… 2026-04-11
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
To address the challenges of model complexity, training inefficiency, and high memory overhead in spatiotemporal forecasting (e.g., traffic flow, meteorological modeling), this paper proposes a lightweight spectral-decoupled knowledge distillation framework. Our method introduces a novel frequency-alignment strategy that decomposes the teacher model’s multi-scale spatiotemporal representations into low- and high-frequency components, which are then selectively transferred to a structurally simplified student network. Specifically, low-frequency components—capturing global trends—are modeled by a convolutional-Transformer hybrid encoder–latent evolution–decoder architecture, while high-frequency components—encoding local dynamics—are captured via spectral analysis. Evaluated on the Navier–Stokes dataset, our approach reduces MSE and MAE by 81.3% and 52.3%, respectively, achieving substantial computational efficiency gains without sacrificing prediction accuracy. This work establishes an optimal trade-off between model fidelity and inference efficiency in spatiotemporal forecasting.

Technology Category

Planning, Routing, and Scheduling: Optimization of Spatio-temporal SystemsData Mining & Knowledge Management: Mining of Spatial, Temporal or Spatio-Temporal DataMachine Learning: Time-Series/Data Streams

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
šŸ“ Abstract
Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation (termed SDKD), which transfers the multi-scale spatiotemporal representations from a complex teacher model to a more efficient lightweight student network. The teacher model follows an encoder-latent evolution-decoder architecture, where its latent evolution module decouples high-frequency details and low-frequency trends using convolution and Transformer (global low-frequency modeler). However, the multi-layer convolution and deconvolution structures result in slow training and high memory usage. To address these issues, we propose a frequency-aligned knowledge distillation strategy, which extracts multi-scale spectral features from the teacher's latent space, including both high and low frequency components, to guide the lightweight student model in capturing both local fine-grained variations and global evolution patterns. Experimental results show that SDKD significantly improves performance, achieving reductions of up to 81.3% in MSE and in MAE 52.3% on the Navier-Stokes equation dataset. The framework effectively captures both high-frequency variations and long-term trends while reducing computational complexity. Our codes are available at https://github.com/itsnotacie/SDKD
Problem

Research questions and friction points this paper is trying to address.

Reducing training inefficiency in spatiotemporal forecasting models
Minimizing high memory consumption in complex forecasting architectures
Transferring multi-scale spectral features to lightweight student models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Frequency-aligned knowledge distillation for efficiency
Multi-scale spectral feature extraction from teacher
Lightweight student captures local and global patterns
šŸ”Ž Similar Papers
No similar papers found.