Towards Robust Multimodal Physiological Foundation Models: Handling Arbitrary Missing Modalities

📅 2025-04-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing methods for modeling multimodal physiological signals (e.g., EEG, ECG) suffer from poor cross-dataset generalization and lack robustness to arbitrary modality missing during inference. To address these challenges, we propose PhysioOmni—the first robust multimodal foundation model specifically designed for physiological signals. Its key innovations include: (1) a decoupled multimodal tokenizer that jointly learns modality-specific and modality-invariant representations; (2) joint masked pretraining with explicit modality-invariance constraints; and (3) a prototype-alignment-based fine-tuning mechanism enabling universal representation transfer under arbitrary subset modality missing. Evaluated on four benchmark tasks—emotion recognition, sleep staging, motion prediction, and mental workload detection—PhysioOmni achieves state-of-the-art performance. Crucially, it demonstrates significant improvements in robustness and cross-dataset generalization under scenarios with 1–3 missing modalities.

Technology Category

Machine Learning: Multimodal LearningComputer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor Fusion

Application Category

User Modeling, Personalization and Recommendation: On-Device user modeling, personalization, and recommendationSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Multimodal physiological signals, such as EEG, ECG, EOG, and EMG, are crucial for healthcare and brain-computer interfaces. While existing methods rely on specialized architectures and dataset-specific fusion strategies, they struggle to learn universal representations that generalize across datasets and handle missing modalities at inference time. To address these issues, we propose PhysioOmni, a foundation model for multimodal physiological signal analysis that models both homogeneous and heterogeneous features to decouple multimodal signals and extract generic representations while maintaining compatibility with arbitrary missing modalities. PhysioOmni trains a decoupled multimodal tokenizer, enabling masked signal pre-training via modality-invariant and modality-specific objectives. To ensure adaptability to diverse and incomplete modality combinations, the pre-trained encoders undergo resilient fine-tuning with prototype alignment on downstream datasets. Extensive experiments on four downstream tasks, emotion recognition, sleep stage classification, motor prediction, and mental workload detection, demonstrate that PhysioOmni achieves state-of-the-art performance while maintaining strong robustness to missing modalities. Our code and model weights will be released.
Problem

Research questions and friction points this paper is trying to address.

Handling arbitrary missing modalities in multimodal physiological signals
Learning universal representations for diverse datasets and tasks
Ensuring robustness and adaptability to incomplete modality combinations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decouples multimodal signals for generic representations
Uses masked signal pre-training with dual objectives
Employs resilient fine-tuning for missing modalities
🔎 Similar Papers
No similar papers found.
X
Xi Fu
Nanyang Technological University
W
Wei-Bang Jiang
Nanyang Technological University, Shanghai Jiao Tong University
Y
Yi Ding
Nanyang Technological University
Cuntai Guan
Cuntai Guan
President's Chair Professor, CCDS, Nanyang Technological University
Brain-Computer InterfaceBrain-Computer InterfacesMachine LearningArtificial Intelligence