🤖 AI Summary
To address label noise in EEG signals—arising from inherent physiological noise and annotation errors—and the modality gap and soft-label misalignment in cross-modal knowledge distillation, this paper proposes an uncertainty-aware prototypical learning framework. The framework integrates a prototype-guided feature semantic alignment module with a task-adaptive distillation head to achieve robust semantic alignment between EEG and visual modalities and noise-resilient knowledge transfer. Evaluated on public multimodal datasets, our method significantly outperforms unimodal baselines and state-of-the-art cross-modal distillation approaches in both emotion classification and regression tasks. It is the first work to jointly unify uncertainty modeling, prototypical learning, and modality-cooperative distillation for EEG-based cognitive state decoding, establishing a principled approach to handling label ambiguity and cross-modal heterogeneity in brain-computer interface applications.
📝 Abstract
Electroencephalography (EEG) is a fundamental modality for cognitive state monitoring in brain-computer interfaces (BCIs). However, it is highly susceptible to intrinsic signal errors and human-induced labeling errors, which lead to label noise and ultimately degrade model performance. To enhance EEG learning, multimodal knowledge distillation (KD) has been explored to transfer knowledge from visual models with rich representations to EEG-based models. Nevertheless, KD faces two key challenges: modality gap and soft label misalignment. The former arises from the heterogeneous nature of EEG and visual feature spaces, while the latter stems from label inconsistencies that create discrepancies between ground truth labels and distillation targets. This paper addresses semantic uncertainty caused by ambiguous features and weakly defined labels. We propose a novel cross-modal knowledge distillation framework that mitigates both modality and label inconsistencies. It aligns feature semantics through a prototype-based similarity module and introduces a task-specific distillation head to resolve label-induced inconsistency in supervision. Experimental results demonstrate that our approach improves EEG-based emotion regression and classification performance, outperforming both unimodal and multimodal baselines on a public multimodal dataset. These findings highlight the potential of our framework for BCI applications.