Score
Design and implement teacher–student training pipelines, loss functions, and mapping modules that transfer representations and semantic knowledge across modalities (e.g., audio, text, vision), including per-layer, multi-layer or hierarchical mappings and spectral/frequency-domain alignment. Build training protocols for multi-to-single, multimodal-to-unimodal or unimodal-to-multimodal distillation that provide stable supervisory signals on unlabeled targets, regularize representation learning, align modality-invariant semantics, and enable students to operate when a teacher modality is absent at inference.
To address insufficient exploitation of complementary prior knowledge, suboptimal distillation path selection, and knowledge drift—arising from data and statistical heterogeneity in cross-modal knowledge distillation—this paper proposes a dynamic adaptive distillation framework. Methodologically, it introduces (1) an instance-level routing network that automatically selects the optimal teacher modality combination per sample, and (2) a plug-and-play masking module, trained independently to suppress modality-specific discrepancies and reconstruct teacher representations. The framework enables flexible integration of both cross-modal and multi-modal teacher models. Evaluated on five benchmark datasets spanning vision, audio, and text modalities, it consistently outperforms state-of-the-art methods, demonstrating superior knowledge transfer efficiency and representation consistency.
Large audio-language models exhibit limited performance on complex reasoning tasks, primarily due to the audio–text modality gap and the absence of structured intermediate supervision. To address this, we propose a unified knowledge distillation framework that transfers symbolic reasoning capabilities from a large text-based teacher model to an audio-based student model while preserving its acoustic understanding. Our approach introduces dual-dimensional distillation—across source modalities (text and audio teachers) and across hierarchical model layers—enabling fine-grained, layer-aligned knowledge transfer. Crucially, we incorporate structured intermediate supervision signals to bridge semantic discrepancies between acoustic representations and symbolic reasoning. Experiments demonstrate substantial improvements in multi-step reasoning performance for audio models, achieving state-of-the-art results across multiple benchmarks and effectively narrowing the semantic gap between speech representation and symbolic reasoning.
This study addresses the challenge that structural heterogeneity in cross-modal feature representations precludes unit-level alignment in conventional knowledge distillation. To overcome this limitation, the authors propose an abstraction mechanism based on vector-quantized codebooks, which transforms teacher features into concept-level anchors to guide student learning. By introducing a task-relevance and compatibility-guided codebook selection algorithm, the proposed method circumvents the assumption of structural consistency across feature spaces, thereby enabling concept-level distillation between heterogeneous modalities without requiring direct alignment. The effectiveness of this framework is validated across multiple classification and semantic segmentation tasks, demonstrating substantial improvements in cross-modal knowledge transfer performance under heterogeneous feature conditions.
To address overfitting in knowledge distillation caused by modality heterogeneity in multimodal learning, this paper proposes a teacher-student framework tailored for discriminative cross-modal knowledge transfer. Methodologically, it integrates cross-modal knowledge distillation, joint feature-classifier alignment, and dynamic sample weighting—without requiring strong inter-modal alignment assumptions. Its key contributions are: (1) a two-level soft-constraint distillation strategy that jointly aligns heterogeneous modalities in both feature space and classifier output space; and (2) a data-quality-aware adaptive sample weighting mechanism to enhance model robustness. Evaluated on speaker identification and image classification tasks, the method significantly improves cross-modal knowledge transfer efficiency and generalization across vision, language, and speech modalities. Notably, it demonstrates superior robustness on low-quality samples, validating its effectiveness under realistic, noisy conditions.
Existing surveys predominantly examine isolated components of multimodal pipelines and lack empirically grounded, pedagogically oriented integration frameworks for teaching and learning contexts. Method: This study introduces the first taxonomy and analytical framework covering five core modalities—natural language, video, sensor data, human-centered signals, and environmental logs—and proposes a novel “mid-fusion” paradigm for multimodal data integration. It further innovates by applying citation graph pruning to achieve structured, high-precision literature synthesis. Contribution/Results: Through systematic review, taxonomic modeling, and multimodal fusion design, we demonstrate that multimodal synergy enables detection of fine-grained learning behaviors imperceptible to unimodal analysis. While prediction accuracy remains largely unchanged, interpretability improves significantly, yielding deeper insights into learners’ cognitive-affective states and training outcomes.
This study addresses the challenge in multi-teacher knowledge distillation where base model preferences interfere with post-training signals, thereby constraining policy composition and routing performance. To overcome this, we propose the Δ-MOPD framework, which introduces a novel objective construction mechanism based on relative teacher shifts. By transferring log-probability shifts and anchoring student initialization, this approach effectively decouples base preferences from post-training increments, establishing objective construction as an independent design dimension. Experimental results demonstrate that under a three-teacher configuration, the method yields a 4.11-point improvement on mathematical tasks and an average gain of 1.95 points across five benchmarks. Furthermore, staged routing significantly reduces sequential discrepancies to 6.42 points. These findings validate the efficacy of the proposed framework for multi-teacher online distillation.
Unified multimodal models often underperform specialized counterparts due to imbalanced learning across modalities. This work proposes a modality-aware teacher routing mechanism that directs student outputs to modality-specific teachers and introduces token-level confidence filtering to distill only the knowledge for which teachers exhibit higher confidence than the student. Additionally, distillation weights are independently tuned per modality, and the model is jointly optimized for both answer correctness and reasoning plausibility. Evaluated across 12 benchmarks and three model scales, the approach achieves state-of-the-art average performance; notably, the 30B variant surpasses all baselines and jointly post-trained models, ranking among the top two on 11 tasks. These results demonstrate that a single general-purpose model can deliver highly effective multimodal capabilities without requiring multiple specialized systems.
This work challenges the common practice in knowledge distillation of naively matching a teacher model’s absolute feature representations, which overlooks the fact that such representations are only equivalent up to orthogonal transformations and isotropic scaling. From a geometric perspective, the paper proposes a new paradigm centered on representation equivalence classes: the student should instead learn class-invariant structures of the teacher’s representations—such as Gram matrices, centered kernel alignment (CKA), or principal subspaces—or leverage coordinate alignment for effective supervision. This framework unifies feature matching, relational distillation, and grafting approaches, revealing that logit-level matching is ultimately key to capability transfer. Experiments on Qwen2.5 and Llama-3.1 demonstrate that high CKA similarity alone is insufficient for performance recovery, while successful grafting hinges on boundary overlap in the training data coverage, thereby validating the proposed theory.
This study addresses the limitations of existing multi-teacher knowledge distillation approaches, which rely on prompt-level domain labels and overlook cross-domain complementary signals, thereby constraining model generalization. To overcome these issues, this work proposes a label-free, token-level routing framework integrated with multi-teacher online policy distillation. Specifically, an expert alignment scoring mechanism is introduced to dynamically weight the supervisory signals from individual teachers, effectively mining and leveraging cross-domain complementary knowledge. By eliminating the dependency on domain labels, the proposed framework achieves substantial performance improvements over existing baselines in both labeled and unlabeled settings, yielding overall score gains of up to 12.3%.
This study addresses the bottleneck in multimodal online policy distillation, where limited student perceptual capacity constrains the performance ceiling achievable through teacher supervision alone. To overcome this limitation, we propose the S-OPD framework, which explicitly enhances student visual understanding via teacher-calibrated strategy contrast and a policy consistency objective. Furthermore, the framework strengthens perceptual learning on the student side by integrating a token-level gating mechanism with policy alignment under image masking and noise perturbation. Notably, S-OPD can be seamlessly incorporated into existing methods without requiring additional data or parameters. Extensive experiments demonstrate significant performance improvements across eight benchmarks, achieving gains of up to 4.25 points on LogicVista. Moreover, the proposed approach yields complementary benefits when combined with teacher-side optimization techniques, establishing it as an effective and versatile solution for enhancing multimodal policy distillation.