design lightweight multimodal models

Designs and implements modular, compute- and parameter-efficient multimodal models that combine visual (RGB) inputs and physiological time-series (e.g., heart-rate, EDA) using lightweight, interchangeable modality modules and efficient fusion layers. These systems are built to run with low latency on resource-constrained hardware, support flexible module composition and missing-modality operation, and enable real-time inference of states derived from combined modalities.

designlightweightmultimodalmodels

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

To address the challenges of label scarcity and modality heterogeneity in multimodal physiological signal fusion (e.g., ECG/EEG), this paper proposes a lightweight, efficient cross-modal learning paradigm. We design a symmetric dual-encoder architecture and introduce a dual masking strategy to enhance self-supervised pretraining in CBraMod. Instead of employing complex fusion modules, we adopt embedding-level concatenation for minimalistic, computationally efficient fusion. Under extremely limited multimodal supervision, our approach achieves performance competitive with state-of-the-art methods on emotion recognition, significantly outperforming conventional multimodal fusion models. The core contribution lies in empirically validating the effectiveness and generalizability of the “foundation model + lightweight fusion” paradigm—demonstrating its scalability and low computational overhead for few-shot multimodal physiological analysis. This work establishes a practical, resource-efficient framework for real-world deployment in data-scarce biomedical scenarios.

Addresses limited labeled data and modality differences in multi-modal physiological signal integration.Develops foundational models for ECG and EEG to capture intra- and inter-modal dependencies.Enables effective downstream learning with simple fusion for healthcare and affective computing.

This work addresses the limitations of existing foundation models for physiological signals, which are often confined to single modalities and struggle with efficient, robust multimodal fusion and edge deployment under data-scarce conditions. The authors propose a lightweight multimodal foundation model with only 5.4 million parameters that enables early cross-modal fusion of EEG, ECG, and PPG through a shared encoder augmented with sensor-type embeddings and a unified query mechanism, offering robustness to missing modalities during inference. Integrated with a channel-unification module, query-set augmentation, INT8 quantization-aware training, and RISC-V deployment optimizations, the model achieves a balanced accuracy of 81.21% on TUAB abnormal EEG detection and a state-of-the-art performance of 0.7416 on HMC sleep staging, while enabling real-time inference on a GAP9 chip with only 18.8 mJ energy consumption and 325.6 ms latency.

cross-modal fusionedge intelligencemissing modalities

Towards Robust Multimodal Physiological Foundation Models: Handling Arbitrary Missing Modalities

Apr 28, 2025
XF
Xi Fu
🏛️ Nanyang Technological University | Shanghai Jiao Tong University

Existing methods for modeling multimodal physiological signals (e.g., EEG, ECG) suffer from poor cross-dataset generalization and lack robustness to arbitrary modality missing during inference. To address these challenges, we propose PhysioOmni—the first robust multimodal foundation model specifically designed for physiological signals. Its key innovations include: (1) a decoupled multimodal tokenizer that jointly learns modality-specific and modality-invariant representations; (2) joint masked pretraining with explicit modality-invariance constraints; and (3) a prototype-alignment-based fine-tuning mechanism enabling universal representation transfer under arbitrary subset modality missing. Evaluated on four benchmark tasks—emotion recognition, sleep staging, motion prediction, and mental workload detection—PhysioOmni achieves state-of-the-art performance. Crucially, it demonstrates significant improvements in robustness and cross-dataset generalization under scenarios with 1–3 missing modalities.

Ensuring robustness and adaptability to incomplete modality combinationsHandling arbitrary missing modalities in multimodal physiological signalsLearning universal representations for diverse datasets and tasks

Current approaches to biosignal time-series classification lack a unified modeling paradigm, limiting both generalization and interpretability. This work proposes a Morphology–Modality Unified Framework, which systematically demonstrates for the first time that waveform morphology—such as spikes, oscillations, and rhythms—rather than model architecture, is the key determinant of performance. The framework integrates morphological features into preprocessing, deep architecture design, and multimodal analysis across EEG, EMG, ECG, and other biosignals. It reveals that the success of deep models stems from the alignment between their inductive biases and the dynamic structure of waveforms. Building on this insight, the study introduces morphology-aware data augmentation strategies and evaluation metrics, substantially enhancing cross-modal generalization and interpretability, thereby establishing morphology-driven modeling as a universal principle in biosignal analysis.

biological signalsmodalitymorphology

Current biomedical multimodal models face critical bottlenecks: reliance on end-to-end training, exponential growth in computational complexity with modality count, severe performance degradation under extreme modality imbalance, and rigid topological coupling. To address these, we propose MM-Lego—a tuning-free, universal multimodal fusion framework. It introduces a novel frequency-domain feature harmonization mechanism that achieves shape alignment and interference-free merging of arbitrary unimodal encoders. We further design modality-agnostic wrappers and zero-/few-shot model merging strategies, enabling topology-agnostic fusion and robust modeling under highly imbalanced modalities. Crucially, MM-Lego requires no fine-tuning yet matches or surpasses end-to-end models in performance, while maintaining full encoder compatibility. Evaluated across seven biomedical benchmark datasets, it achieves state-of-the-art results on five—demonstrating unprecedented flexibility, efficiency, and generalizability in biomedical multimodal learning.

Enabling flexible model merging with minimal fine-tuningHandling diverse data modalities in biomedical machine learningOvercoming limitations of existing multimodal fusion approaches

Latest Papers

What's happening recently
View more

Existing self-supervised approaches treat multi-site physiological signals as interchangeable views, overlooking their intrinsic temporal dynamics and physiological conduction relationships. This work proposes xMAE, a novel framework that, for the first time, incorporates directional temporal structure into multimodal self-supervised pretraining. By leveraging a masked cross-modal reconstruction mechanism, xMAE explicitly models the physiological constraints and temporal delays between signals such as ECG and PPG to capture realistic conduction processes. Evaluated across 19 downstream tasks—including cardiovascular prediction, anomaly detection, and sleep staging—the method outperforms both unimodal and multimodal baselines on 15 tasks and demonstrates strong generalization across diverse devices and acquisition conditions.

biosignal representation learningcross-modal reconstructionphysiological process

MedM2T: A MultiModal Framework for Time-Aware Modeling with Electronic Health Record and Electrocardiogram Data

Oct 31, 2025
YK
Yu-Chen Kuo
🏛️ National Yang Ming Chiao Tung University | Boston Children’s Hospital

To address the temporal heterogeneity, sparsity, and modality granularity discrepancies inherent in multimodal clinical time-series data—such as electronic health records (EHR) and electrocardiograms (ECG)—this paper proposes a hierarchical multimodal temporal fusion model. Methodologically, it introduces a sparse time-series encoder, a hierarchical temporal fusion module, and a dual-modal attention mechanism, augmented by modality-specific pretrained encoders and a shared latent-space feature alignment strategy to enable dynamic cross-modal interaction and unified multi-granularity temporal representation learning. Evaluated on MIMIC-IV and MIMIC-IV-ECG, the model achieves state-of-the-art performance: AUROC = 0.947 for 90-day cardiovascular event prediction, AUROC = 0.901 for in-hospital mortality prediction, and MAE = 2.31 hours for ICU length-of-stay regression. It demonstrates strong generalizability and scalability across diverse clinical forecasting tasks.

Handling multimodal medical data with complex temporal structuresIntegrating sparse EHR and dense ECG time series dataPredicting chronic and acute disease outcomes clinically

Existing clinical multimodal AI systems lack systematic evaluation of the robustness of fusion architectures under sensor failure scenarios, including both complete modality absence and continuous intra-modality missingness. To address this gap, this work proposes MuteBench, a comprehensive benchmark spanning seven clinical domains, nine datasets, six fusion architectures, and two missingness patterns, enabling unified assessment of model fault tolerance under controlled levels of missing data severity. The study reveals that architecture type is the strongest predictor of robustness; channel-independent models exhibit resilience to full modality dropout but heightened sensitivity to intra-modality gaps; curriculum-based modality dropout proves effective only within the maximum dropout rate observed during training; and diffusion-based imputation substantially enhances classification performance for models otherwise vulnerable to input corruption.

clinical AImodality missingmultimodal fusion

This work addresses the challenge in multimodal time series forecasting where unconstrained fusion of auxiliary modalities—such as text or vision—often introduces irrelevant information that degrades predictive performance. To mitigate this issue, the authors propose the Controllable Fusion Adapter (CFA), a plug-and-play, low-rank adaptation module that selectively integrates only those cross-modal features aligned with temporal dynamics, without modifying the backbone model. CFA is compatible with diverse time series and text architectures and employs a controlled fusion mechanism to efficiently filter relevant information. Extensive evaluation across multiple datasets, encompassing over 20,000 experiments, demonstrates that CFA consistently outperforms existing fusion strategies, underscoring its effectiveness and generalizability.

auxiliary modalitiesconstrained fusionirrelevant information

This work addresses the limitations of existing multimodal fusion approaches for electrocardiogram (ECG) and clinical text, which often neglect the spatiotemporal dependencies among ECG leads and are susceptible to modality bias from textual inputs, leading to inaccurate diagnostic representations. To overcome these issues, the authors propose a decoupled multimodal ECG representation learning framework that captures fine-grained dynamic features through spatiotemporal masked modeling. The framework integrates contrastive learning with a generative reconstruction mechanism and employs both modality-shared and modality-specific encoders to effectively disentangle modality-invariant and modality-specific information. Extensive experiments on three public datasets demonstrate significant performance gains on downstream tasks, underscoring the method’s advantages in robustness, interpretability, and generalization capability.

inter-modalityintra-modalitymodality-specific bias

Hot Scholars

LC

Lequn Chen

University of Washington
Distributed SystemsMachine Learning SystemsOperating Systems
SK

Seung Ki Moon

Associate Professor of Mechanical and Aerospace Engineering, Nanyang Technological University
Product family and platform designadditive manufacturing and 3D printingoptimizationdigital twins
ZZ

Zichen Zhu

Shanghai Jiao Tong University
GUI智能体,多模态大模型,人机交互
ZZ

Zihan Zhao

Shanghai Jiao Tong University
NLP
YL

Yansi Li

Shanghai Jiao Tong University
Large Language ModelsReasoningGUI Agents