Leveraging Foundational Models and Simple Fusion for Multi-modal Physiological Signal Analysis

📅 2025-12-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of label scarcity and modality heterogeneity in multimodal physiological signal fusion (e.g., ECG/EEG), this paper proposes a lightweight, efficient cross-modal learning paradigm. We design a symmetric dual-encoder architecture and introduce a dual masking strategy to enhance self-supervised pretraining in CBraMod. Instead of employing complex fusion modules, we adopt embedding-level concatenation for minimalistic, computationally efficient fusion. Under extremely limited multimodal supervision, our approach achieves performance competitive with state-of-the-art methods on emotion recognition, significantly outperforming conventional multimodal fusion models. The core contribution lies in empirically validating the effectiveness and generalizability of the “foundation model + lightweight fusion” paradigm—demonstrating its scalability and low computational overhead for few-shot multimodal physiological analysis. This work establishes a practical, resource-efficient framework for real-world deployment in data-scarce biomedical scenarios.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Physiological signals such as electrocardiograms (ECG) and electroencephalograms (EEG) provide complementary insights into human health and cognition, yet multi-modal integration is challenging due to limited multi-modal labeled data, and modality-specific differences . In this work, we adapt the CBraMod encoder for large-scale self-supervised ECG pretraining, introducing a dual-masking strategy to capture intra- and inter-lead dependencies. To overcome the above challenges, we utilize a pre-trained CBraMod encoder for EEG and pre-train a symmetric ECG encoder, equipping each modality with a rich foundational representation. These representations are then fused via simple embedding concatenation, allowing the classification head to learn cross-modal interactions, together enabling effective downstream learning despite limited multi-modal supervision. Evaluated on emotion recognition, our approach achieves near state-of-the-art performance, demonstrating that carefully designed physiological encoders, even with straightforward fusion, substantially improve downstream performance. These results highlight the potential of foundation-model approaches to harness the holistic nature of physiological signals, enabling scalable, label-efficient, and generalizable solutions for healthcare and affective computing.
Problem

Research questions and friction points this paper is trying to address.

Addresses limited labeled data and modality differences in multi-modal physiological signal integration.
Develops foundational models for ECG and EEG to capture intra- and inter-modal dependencies.
Enables effective downstream learning with simple fusion for healthcare and affective computing.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-supervised pretraining for ECG using dual-masking strategy
Utilizing pre-trained foundational models for both ECG and EEG
Simple concatenation fusion enabling cross-modal interaction learning
🔎 Similar Papers
No similar papers found.
Alexandria University
Y
Youssef Ghallab
Computer and Communication Engineering Department, Alexandria University
O
Omar Iraqy
Computer and Communication Engineering Department, Alexandria University
Mohamed Kandil
Mohamed Kandil
Computer and Communication Engineering Department, Alexandria University
M
Mohamed Ashraf
Computer and Communication Engineering Department, Alexandria University
S
Saadeldine Eletter
Computer and Communication Engineering Department, Alexandria University
M
Morougue Ghazal
Computer and Communication Engineering Department, Alexandria University
A
Ayman Khalafallah
Computer and Communication Engineering Department, Alexandria University
Nagwa El-Makky
Nagwa El-Makky
Computer and Communication Engineering Department, Alexandria University