cross-modal knowledge distillation

Design and implement teacher–student training pipelines, loss functions, and mapping modules that transfer representations and semantic knowledge across modalities (e.g., audio, text, vision), including per-layer, multi-layer or hierarchical mappings and spectral/frequency-domain alignment. Build training protocols for multi-to-single, multimodal-to-unimodal or unimodal-to-multimodal distillation that provide stable supervisory signals on unlabeled targets, regularize representation learning, align modality-invariant semantics, and enable students to operate when a teacher modality is absent at inference.

cross-modalknowledgedistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address insufficient exploitation of complementary prior knowledge, suboptimal distillation path selection, and knowledge drift—arising from data and statistical heterogeneity in cross-modal knowledge distillation—this paper proposes a dynamic adaptive distillation framework. Methodologically, it introduces (1) an instance-level routing network that automatically selects the optimal teacher modality combination per sample, and (2) a plug-and-play masking module, trained independently to suppress modality-specific discrepancies and reconstruct teacher representations. The framework enables flexible integration of both cross-modal and multi-modal teacher models. Evaluated on five benchmark datasets spanning vision, audio, and text modalities, it consistently outperforms state-of-the-art methods, demonstrating superior knowledge transfer efficiency and representation consistency.

Addresses cross-modal knowledge distillation challengesEnhances transfer with specialized teacher mixturesSolves distillation path selection and knowledge drift

Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation

Sep 22, 2025
RY
Runyan Yang
🏛️ Jiutian Artificial Intelligence Research Institute, China Mobile | The State Key Laboratory of Multimedia Information Processing, Peking University

Large audio-language models exhibit limited performance on complex reasoning tasks, primarily due to the audio–text modality gap and the absence of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework that transfers symbolic reasoning capabilities from a large text-based teacher model to an audio-based student model while preserving its acoustic understanding. Our approach introduces dual-dimensional distillation—across source modalities (text and audio teachers) and across hierarchical model layers—enabling fine-grained, layer-aligned knowledge transfer. Crucially, we incorporate structured intermediate supervision signals to bridge semantic discrepancies between acoustic representations and symbolic reasoning. Experiments demonstrate substantial improvements in multi-step reasoning performance for audio models, achieving state-of-the-art results across multiple benchmarks and effectively narrowing the semantic gap between speech representation and symbolic reasoning.

Addressing lack of structured intermediate supervision in audio modelsBridging the modality gap between audio and text representationsTransferring reasoning capabilities from text to audio models

This study addresses the challenge that structural heterogeneity in cross-modal feature representations precludes unit-level alignment in conventional knowledge distillation. To overcome this limitation, the authors propose an abstraction mechanism based on vector-quantized codebooks, which transforms teacher features into concept-level anchors to guide student learning. By introducing a task-relevance and compatibility-guided codebook selection algorithm, the proposed method circumvents the assumption of structural consistency across feature spaces, thereby enabling concept-level distillation between heterogeneous modalities without requiring direct alignment. The effectiveness of this framework is validated across multiple classification and semantic segmentation tasks, demonstrating substantial improvements in cross-modal knowledge transfer performance under heterogeneous feature conditions.

Cross-modal knowledge distillationFeature-level alignmentStructurally heterogeneous features

Cross-Modal Distillation For Widely Differing Modalities

Jul 22, 2025
CZ
Cairong Zhao
🏛️ Tongji University | Alibaba Group | Oosto

To address overfitting in knowledge distillation caused by modality heterogeneity in multimodal learning, this paper proposes a teacher-student framework tailored for discriminative cross-modal knowledge transfer. Methodologically, it integrates cross-modal knowledge distillation, joint feature-classifier alignment, and dynamic sample weighting—without requiring strong inter-modal alignment assumptions. Its key contributions are: (1) a two-level soft-constraint distillation strategy that jointly aligns heterogeneous modalities in both feature space and classifier output space; and (2) a data-quality-aware adaptive sample weighting mechanism to enhance model robustness. Evaluated on speaker identification and image classification tasks, the method significantly improves cross-modal knowledge transfer efficiency and generalization across vision, language, and speech modalities. Notably, it demonstrates superior robustness on low-quality samples, validating its effectiveness under realistic, noisy conditions.

Improve robustness via quality-based adaptive weightsPrevent overfitting in cross-modal distillationTransfer knowledge between widely differing modalities

Existing surveys predominantly examine isolated components of multimodal pipelines and lack empirically grounded, pedagogically oriented integration frameworks for teaching and learning contexts. Method: This study introduces the first taxonomy and analytical framework covering five core modalities—natural language, video, sensor data, human-centered signals, and environmental logs—and proposes a novel “mid-fusion” paradigm for multimodal data integration. It further innovates by applying citation graph pruning to achieve structured, high-precision literature synthesis. Contribution/Results: Through systematic review, taxonomic modeling, and multimodal fusion design, we demonstrate that multimodal synergy enables detection of fine-grained learning behaviors imperceptible to unimodal analysis. While prediction accuracy remains largely unchanged, interpretability improves significantly, yielding deeper insights into learners’ cognitive-affective states and training outcomes.

Addresses challenges in real-time multimodal data integrationIntroduces taxonomy for five modality groups and data fusionReviews empirical multimodal methods in learning environments

Latest Papers

What's happening recently
View more

This study addresses the challenge in multi-teacher knowledge distillation where base model preferences interfere with post-training signals, thereby constraining policy composition and routing performance. To overcome this, we propose the Δ-MOPD framework, which introduces a novel objective construction mechanism based on relative teacher shifts. By transferring log-probability shifts and anchoring student initialization, this approach effectively decouples base preferences from post-training increments, establishing objective construction as an independent design dimension. Experimental results demonstrate that under a three-teacher configuration, the method yields a 4.11-point improvement on mathematical tasks and an average gain of 1.95 points across five benchmarks. Furthermore, staged routing significantly reduces sequential discrepancies to 6.42 points. These findings validate the efficacy of the proposed framework for multi-teacher online distillation.

endpoint policyknowledge transferlogit shift

Unified multimodal models often underperform specialized counterparts due to imbalanced learning across modalities. This work proposes a modality-aware teacher routing mechanism that directs student outputs to modality-specific teachers and introduces token-level confidence filtering to distill only the knowledge for which teachers exhibit higher confidence than the student. Additionally, distillation weights are independently tuned per modality, and the model is jointly optimized for both answer correctness and reasoning plausibility. Evaluated across 12 benchmarks and three model scales, the approach achieves state-of-the-art average performance; notably, the 30B variant surpasses all baselines and jointly post-trained models, ranking among the top two on 11 tasks. These results demonstrate that a single general-purpose model can deliver highly effective multimodal capabilities without requiring multiple specialized systems.

cross-modal interferencemodality balancemultimodal learning

This work challenges the common practice in knowledge distillation of naively matching a teacher model’s absolute feature representations, which overlooks the fact that such representations are only equivalent up to orthogonal transformations and isotropic scaling. From a geometric perspective, the paper proposes a new paradigm centered on representation equivalence classes: the student should instead learn class-invariant structures of the teacher’s representations—such as Gram matrices, centered kernel alignment (CKA), or principal subspaces—or leverage coordinate alignment for effective supervision. This framework unifies feature matching, relational distillation, and grafting approaches, revealing that logit-level matching is ultimately key to capability transfer. Experiments on Qwen2.5 and Llama-3.1 demonstrate that high CKA similarity alone is insufficient for performance recovery, while successful grafting hinges on boundary overlap in the training data coverage, thereby validating the proposed theory.

capability recoverygeometric invarianceknowledge distillation

This study addresses the limitations of existing multi-teacher knowledge distillation approaches, which rely on prompt-level domain labels and overlook cross-domain complementary signals, thereby constraining model generalization. To overcome these issues, this work proposes a label-free, token-level routing framework integrated with multi-teacher online policy distillation. Specifically, an expert alignment scoring mechanism is introduced to dynamically weight the supervisory signals from individual teachers, effectively mining and leveraging cross-domain complementary knowledge. By eliminating the dependency on domain labels, the proposed framework achieves substantial performance improvements over existing baselines in both labeled and unlabeled settings, yielding overall score gains of up to 12.3%.

Cross-domain supervisionMulti-teacher on-policy distillationTeacher routing

This study addresses the bottleneck in multimodal online policy distillation, where limited student perceptual capacity constrains the performance ceiling achievable through teacher supervision alone. To overcome this limitation, we propose the S-OPD framework, which explicitly enhances student visual understanding via teacher-calibrated strategy contrast and a policy consistency objective. Furthermore, the framework strengthens perceptual learning on the student side by integrating a token-level gating mechanism with policy alignment under image masking and noise perturbation. Notably, S-OPD can be seamlessly incorporated into existing methods without requiring additional data or parameters. Extensive experiments demonstrate significant performance improvements across eight benchmarks, achieving gains of up to 4.25 points on LogicVista. Moreover, the proposed approach yields complementary benefits when combined with teacher-side optimization techniques, establishing it as an effective and versatile solution for enhancing multimodal policy distillation.

Knowledge DistillationMultimodal On-Policy DistillationStudent Perception

Hot Scholars

XL

Xitong Ling

Tsinghua University
AI4PathologyFoundation-ModelVision-Language-Model
LN

Linh Ngo Van

Hanoi University of Science and Technology
Machine LearningData MiningNatural Language Processing
YH

Yonghong He

清华大学深圳国际研究生院
生物医学工程,光学成像,AI图像处理、病理大模型
TL

Trung Le

Faculty of Information Technology, Monash University, Australia
Adversarial Machine LearningGenerative ModelsModel UnlearningModel Editing
TH

Thien Huu Nguyen

University of Oregon
Information ExtractionDeep LearningNatural Language ProcessingMachine Learning