๐ค AI Summary
This work addresses the degradation in classification performance caused by insufficiently discriminative patterns in short temporal inputs and the difficulty of effectively transferring knowledge from long-context teacher models. To this end, it introduces diffusion priors into the knowledge distillation framework for the first time. The method treats the student modelโs short-context features as degraded observations of the teacherโs long-context representations and leverages a diffusion model to generate diverse, long-context supervisory signals. Task-relevant knowledge is then transferred through Bayesian posterior sampling, enabling distributed and adaptive distillation. Extensive experiments demonstrate that the proposed approach significantly improves short-sequence classification accuracy across various early-exit configurations, datasets, and model architectures, effectively narrowing the generalization gap induced by input length discrepancies.
๐ Abstract
While traditional time-series classifiers assume full sequences at inference, practical constraints (latency and cost) often limit inputs to partial prefixes. The absence of class-discriminative patterns in partial data can significantly hinder a classifier's ability to generalize. This work uses knowledge distillation (KD) to equip partial time series classifiers with the generalization ability of their full-sequence counterparts. In KD, high-capacity teacher transfers supervision to aid student learning on the target task. Matching with teacher features has shown promise in closing the generalization gap due to limited parameter capacity. However, when the generalization gap arises from training-data differences (full versus partial), the teacher's full-context features can be an overwhelming target signal for the student's short-context features. To provide progressive, diverse, and collective teacher supervision, we propose Generative Diffusion Prior Distillation (GDPD), a novel KD framework that treats short-context student features as degraded observations of the target full-context features. Inspired by the iterative restoration capability of diffusion models, we learn a diffusion-based generative prior over teacher features. Leveraging this prior, we posterior-sample target teacher representations that could best explain the missing long-range information in the student features and optimize the student features to be minimally degraded relative to these targets. GDPD provides each student feature with a distribution of task-relevant long-context knowledge, which benefits learning on the partial classification task. Extensive experiments across earliness settings, datasets, and architectures demonstrate GDPD's effectiveness for full-to-partial distillation.