🤖 AI Summary
This study addresses the limitation of existing knowledge distillation methods that overlook cross-sample predictive structures, thereby constraining the efficiency and accuracy of spatiotemporal prediction models. To this end, we propose a task-aware memory distillation framework that introduces a bounded, retrievable memory bank to store historical reference samples from the teacher model. By integrating task-specific selection rules with a residual prototype regression mechanism, the framework enhances the student's capacity to exploit cross-sample structural information. Furthermore, feature matching and lightweight adapters are incorporated to optimize the training process. Extensive experiments across video, weather, and traffic forecasting tasks demonstrate that the proposed method significantly reduces mean squared error while improving structural similarity index scores, achieving notable accuracy gains without introducing additional inference overhead.
📝 Abstract
Knowledge distillation enables efficient spatiotemporal prediction by transferring knowledge from an accurate teacher to a compact student. However, matching outputs or features independently for each sample leaves cross-sample predictive structure underused. Exploiting this structure requires representations and historical references that reflect the dynamics of each task. We propose TAM, a Task-Aware Memory Distillation framework that organizes a frozen teacher's knowledge into a bounded, retrievable history. Memory entries encode latent features, forecast changes, or flow residuals, while task-specific selection rules identify relevant historical references. The student either matches the teacher's similarity distribution over shared references or regresses observation-conditioned residual prototypes. These objectives complement supervised prediction and conventional distillation. The teacher, memory, and auxiliary adapters are used only during training, leaving student inference unchanged. We evaluate TAM on video prediction, weather forecasting, and traffic flow prediction across multiple teacher-student configurations. Averaged over four paired runs, adding TAM improves SSIM on all six video datasets and reduces MSE on five relative to the corresponding KD baselines. Mean paired MSE reductions reach 1.86% on KittiCaltech, 1.93% on WeatherBench with a gSTA teacher, and 1.01% on TaxiBJ. These results demonstrate the utility of historical teacher supervision across distinct forecasting tasks without additional student inference cost.