🤖 AI Summary
This study addresses the challenge that structural heterogeneity in cross-modal feature representations precludes unit-level alignment in conventional knowledge distillation. To overcome this limitation, the authors propose an abstraction mechanism based on vector-quantized codebooks, which transforms teacher features into concept-level anchors to guide student learning. By introducing a task-relevance and compatibility-guided codebook selection algorithm, the proposed method circumvents the assumption of structural consistency across feature spaces, thereby enabling concept-level distillation between heterogeneous modalities without requiring direct alignment. The effectiveness of this framework is validated across multiple classification and semantic segmentation tasks, demonstrating substantial improvements in cross-modal knowledge transfer performance under heterogeneous feature conditions.
📝 Abstract
Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are structurally heterogeneous and lack clear unit-level correspondence, such as 2D spatial visual grids and 1D temporal audio sequences, thereby limiting the applicability of feature-level alignment. To address this challenge, we propose a cross-modal distillation framework that enables effective knowledge transfer across structurally heterogeneous feature spaces via a vector-quantized codebook. Specifically, teacher features are abstracted into a set of vector-form codes regardless of their original feature structure, and the selected codes serve as concept-level anchors for student learning. Code selection is guided by both task relevance and student compatibility, allowing the student to receive transferable teacher knowledge without requiring direct unit-level feature alignment. Experimental results across diverse cross-modal distillation scenarios demonstrate the effectiveness of the proposed framework on classification and semantic segmentation tasks.