Score
Designs and implements modular, compute- and parameter-efficient multimodal models that combine visual (RGB) inputs and physiological time-series (e.g., heart-rate, EDA) using lightweight, interchangeable modality modules and efficient fusion layers. These systems are built to run with low latency on resource-constrained hardware, support flexible module composition and missing-modality operation, and enable real-time inference of states derived from combined modalities.
This work addresses the parameter efficiency challenge in visual–language large models (VLLMs) for effective cross-modal fusion. We systematically analyze 34 state-of-the-art VLLMs and, for the first time, unify their training paradigms into three categories—single-stage fine-tuning, two-stage fine-tuning, and direct adaptation—establishing the first taxonomy of VLLM efficiency grounded in training methodology. Our study fills a critical gap by providing the first systematic analysis of direct adaptation, empirically demonstrating that it achieves over 90% of two-stage fine-tuning performance with less than 1% parameter overhead. We comprehensively examine core components—including LLM backbones, vision encoders, multimodal fusion architectures, parameter-efficient adaptation techniques (e.g., LoRA, Adapters), and evaluation protocols—and synthesize key benchmarks and metrics. The work delivers both a theoretical framework and empirical evidence to advance efficient multimodal modeling.
To address the challenges of label scarcity and modality heterogeneity in multimodal physiological signal fusion (e.g., ECG/EEG), this paper proposes a lightweight, efficient cross-modal learning paradigm. We design a symmetric dual-encoder architecture and introduce a dual masking strategy to enhance self-supervised pretraining in CBraMod. Instead of employing complex fusion modules, we adopt embedding-level concatenation for minimalistic, computationally efficient fusion. Under extremely limited multimodal supervision, our approach achieves performance competitive with state-of-the-art methods on emotion recognition, significantly outperforming conventional multimodal fusion models. The core contribution lies in empirically validating the effectiveness and generalizability of the “foundation model + lightweight fusion” paradigm—demonstrating its scalability and low computational overhead for few-shot multimodal physiological analysis. This work establishes a practical, resource-efficient framework for real-world deployment in data-scarce biomedical scenarios.
This work addresses the limitations of existing foundation models for physiological signals, which are often confined to single modalities and struggle with efficient, robust multimodal fusion and edge deployment under data-scarce conditions. The authors propose a lightweight multimodal foundation model with only 5.4 million parameters that enables early cross-modal fusion of EEG, ECG, and PPG through a shared encoder augmented with sensor-type embeddings and a unified query mechanism, offering robustness to missing modalities during inference. Integrated with a channel-unification module, query-set augmentation, INT8 quantization-aware training, and RISC-V deployment optimizations, the model achieves a balanced accuracy of 81.21% on TUAB abnormal EEG detection and a state-of-the-art performance of 0.7416 on HMC sleep staging, while enabling real-time inference on a GAP9 chip with only 18.8 mJ energy consumption and 325.6 ms latency.
Existing methods for modeling multimodal physiological signals (e.g., EEG, ECG) suffer from poor cross-dataset generalization and lack robustness to arbitrary modality missing during inference. To address these challenges, we propose PhysioOmni—the first robust multimodal foundation model specifically designed for physiological signals. Its key innovations include: (1) a decoupled multimodal tokenizer that jointly learns modality-specific and modality-invariant representations; (2) joint masked pretraining with explicit modality-invariance constraints; and (3) a prototype-alignment-based fine-tuning mechanism enabling universal representation transfer under arbitrary subset modality missing. Evaluated on four benchmark tasks—emotion recognition, sleep staging, motion prediction, and mental workload detection—PhysioOmni achieves state-of-the-art performance. Crucially, it demonstrates significant improvements in robustness and cross-dataset generalization under scenarios with 1–3 missing modalities.
Current approaches to biosignal time-series classification lack a unified modeling paradigm, limiting both generalization and interpretability. This work proposes a Morphology–Modality Unified Framework, which systematically demonstrates for the first time that waveform morphology—such as spikes, oscillations, and rhythms—rather than model architecture, is the key determinant of performance. The framework integrates morphological features into preprocessing, deep architecture design, and multimodal analysis across EEG, EMG, ECG, and other biosignals. It reveals that the success of deep models stems from the alignment between their inductive biases and the dynamic structure of waveforms. Building on this insight, the study introduces morphology-aware data augmentation strategies and evaluation metrics, substantially enhancing cross-modal generalization and interpretability, thereby establishing morphology-driven modeling as a universal principle in biosignal analysis.
Current biomedical multimodal models face critical bottlenecks: reliance on end-to-end training, exponential growth in computational complexity with modality count, severe performance degradation under extreme modality imbalance, and rigid topological coupling. To address these, we propose MM-Lego—a tuning-free, universal multimodal fusion framework. It introduces a novel frequency-domain feature harmonization mechanism that achieves shape alignment and interference-free merging of arbitrary unimodal encoders. We further design modality-agnostic wrappers and zero-/few-shot model merging strategies, enabling topology-agnostic fusion and robust modeling under highly imbalanced modalities. Crucially, MM-Lego requires no fine-tuning yet matches or surpasses end-to-end models in performance, while maintaining full encoder compatibility. Evaluated across seven biomedical benchmark datasets, it achieves state-of-the-art results on five—demonstrating unprecedented flexibility, efficiency, and generalizability in biomedical multimodal learning.
Existing self-supervised approaches treat multi-site physiological signals as interchangeable views, overlooking their intrinsic temporal dynamics and physiological conduction relationships. This work proposes xMAE, a novel framework that, for the first time, incorporates directional temporal structure into multimodal self-supervised pretraining. By leveraging a masked cross-modal reconstruction mechanism, xMAE explicitly models the physiological constraints and temporal delays between signals such as ECG and PPG to capture realistic conduction processes. Evaluated across 19 downstream tasks—including cardiovascular prediction, anomaly detection, and sleep staging—the method outperforms both unimodal and multimodal baselines on 15 tasks and demonstrates strong generalization across diverse devices and acquisition conditions.
To address the temporal heterogeneity, sparsity, and modality granularity discrepancies inherent in multimodal clinical time-series data—such as electronic health records (EHR) and electrocardiograms (ECG)—this paper proposes a hierarchical multimodal temporal fusion model. Methodologically, it introduces a sparse time-series encoder, a hierarchical temporal fusion module, and a dual-modal attention mechanism, augmented by modality-specific pretrained encoders and a shared latent-space feature alignment strategy to enable dynamic cross-modal interaction and unified multi-granularity temporal representation learning. Evaluated on MIMIC-IV and MIMIC-IV-ECG, the model achieves state-of-the-art performance: AUROC = 0.947 for 90-day cardiovascular event prediction, AUROC = 0.901 for in-hospital mortality prediction, and MAE = 2.31 hours for ICU length-of-stay regression. It demonstrates strong generalizability and scalability across diverse clinical forecasting tasks.
Existing clinical multimodal AI systems lack systematic evaluation of the robustness of fusion architectures under sensor failure scenarios, including both complete modality absence and continuous intra-modality missingness. To address this gap, this work proposes MuteBench, a comprehensive benchmark spanning seven clinical domains, nine datasets, six fusion architectures, and two missingness patterns, enabling unified assessment of model fault tolerance under controlled levels of missing data severity. The study reveals that architecture type is the strongest predictor of robustness; channel-independent models exhibit resilience to full modality dropout but heightened sensitivity to intra-modality gaps; curriculum-based modality dropout proves effective only within the maximum dropout rate observed during training; and diffusion-based imputation substantially enhances classification performance for models otherwise vulnerable to input corruption.
This work addresses the challenge in multimodal time series forecasting where unconstrained fusion of auxiliary modalities—such as text or vision—often introduces irrelevant information that degrades predictive performance. To mitigate this issue, the authors propose the Controllable Fusion Adapter (CFA), a plug-and-play, low-rank adaptation module that selectively integrates only those cross-modal features aligned with temporal dynamics, without modifying the backbone model. CFA is compatible with diverse time series and text architectures and employs a controlled fusion mechanism to efficiently filter relevant information. Extensive evaluation across multiple datasets, encompassing over 20,000 experiments, demonstrates that CFA consistently outperforms existing fusion strategies, underscoring its effectiveness and generalizability.
This work addresses the limitations of existing multimodal fusion approaches for electrocardiogram (ECG) and clinical text, which often neglect the spatiotemporal dependencies among ECG leads and are susceptible to modality bias from textual inputs, leading to inaccurate diagnostic representations. To overcome these issues, the authors propose a decoupled multimodal ECG representation learning framework that captures fine-grained dynamic features through spatiotemporal masked modeling. The framework integrates contrastive learning with a generative reconstruction mechanism and employs both modality-shared and modality-specific encoders to effectively disentangle modality-invariant and modality-specific information. Extensive experiments on three public datasets demonstrate significant performance gains on downstream tasks, underscoring the method’s advantages in robustness, interpretability, and generalization capability.