🤖 AI Summary
To address the challenges of label scarcity and modality heterogeneity in multimodal physiological signal fusion (e.g., ECG/EEG), this paper proposes a lightweight, efficient cross-modal learning paradigm. We design a symmetric dual-encoder architecture and introduce a dual masking strategy to enhance self-supervised pretraining in CBraMod. Instead of employing complex fusion modules, we adopt embedding-level concatenation for minimalistic, computationally efficient fusion. Under extremely limited multimodal supervision, our approach achieves performance competitive with state-of-the-art methods on emotion recognition, significantly outperforming conventional multimodal fusion models. The core contribution lies in empirically validating the effectiveness and generalizability of the “foundation model + lightweight fusion” paradigm—demonstrating its scalability and low computational overhead for few-shot multimodal physiological analysis. This work establishes a practical, resource-efficient framework for real-world deployment in data-scarce biomedical scenarios.
📝 Abstract
Physiological signals such as electrocardiograms (ECG) and electroencephalograms (EEG) provide complementary insights into human health and cognition, yet multi-modal integration is challenging due to limited multi-modal labeled data, and modality-specific differences . In this work, we adapt the CBraMod encoder for large-scale self-supervised ECG pretraining, introducing a dual-masking strategy to capture intra- and inter-lead dependencies. To overcome the above challenges, we utilize a pre-trained CBraMod encoder for EEG and pre-train a symmetric ECG encoder, equipping each modality with a rich foundational representation. These representations are then fused via simple embedding concatenation, allowing the classification head to learn cross-modal interactions, together enabling effective downstream learning despite limited multi-modal supervision. Evaluated on emotion recognition, our approach achieves near state-of-the-art performance, demonstrating that carefully designed physiological encoders, even with straightforward fusion, substantially improve downstream performance. These results highlight the potential of foundation-model approaches to harness the holistic nature of physiological signals, enabling scalable, label-efficient, and generalizable solutions for healthcare and affective computing.