🤖 AI Summary
This study addresses the challenges of insufficient spatiotemporal feature extraction and poor cross-subject generalization in perceived speech decoding for non-invasive brain-computer interfaces. To this end, we propose SICMD, a novel method that introduces the first unified framework for fusing multimodal brain signals from fMRI and MEG. By leveraging a deep encoder architecture alongside cross-modal decoding techniques, SICMD synergistically optimizes spatiotemporal representation learning and cross-subject invariance. Experimental results demonstrate that SICMD achieves a cross-subject Top-1 decoding accuracy improvement of over 10.6% compared to baseline methods while reducing training costs by 88.8%, representing a significant breakthrough in both decoding performance and computational efficiency.
📝 Abstract
Perceived speech decoding based on non-invasive brain-computer interface (BCI) signals has been extensively studied in recent years. Research in this field primarily faces two challenges: extracting neural representations with rich spatiotemporal information and achieving cross-subject generalization. Although separate studies have proposed methods to cope with these issues, a unified approach that simultaneously tackles both challenges remains lacking. To fill this gap, we propose the Subject-Invariant Cross-Modal Perceived Speech Decoding (SICMD) method, which integrates functional magnetic resonance imaging (fMRI) and magnetoencephalography (MEG). We conduct comprehensive analyses of the fusion method, fusion position, encoder architecture, and model inputs. Our results demonstrate that the proposed method improves Top-1, Top-10, and Rankacc by more than 10.6%, 10.1%, and 1.7%, respectively, compared to baseline methods in cross-subject perceived speech decoding tasks, while reducing training costs by 88.8% and 60.5% compared to multi-subject and intra-subject decoding settings. Further visualization experiments also confirm the effectiveness of our approach.