🤖 AI Summary
This study addresses the challenge of extending EEG foundation models to multimodal scenarios using unlabeled data while preserving pretrained knowledge. To this end, it proposes ZeroMAG, a framework that introduces a configuration-agnostic adapter generation mechanism. By constructing modality-subject-task conditions, ZeroMAG learns source adapters within a function-constrained latent space to dynamically generate target weights. This enables zero-shot multimodal adaptation of frozen EEG models without requiring target labels or parameter optimization. The proposed method improves balanced accuracy by 7.22% over EEG-only baselines, approaching the performance of supervised multimodal adaptation and validating the critical role of functional supervision in adapter generation.
📝 Abstract
EEG foundation models (EFMs) capture reusable knowledge from large-scale EEG data, while many EEG recordings also include companion physiological signals that provide complementary information beyond the EEG-only interface. The challenge is to preserve this pretrained knowledge while extending the EFM to heterogeneous multimodal recordings through an adaptation inferred from unlabeled target data. We introduce ZeroMAG, a zero-shot multimodal adapter generation framework that extends a frozen EEG encoder and prediction head using unlabeled target recordings, without target labels or target-side optimization. The target datasets are held out from all model training and selection in the ZeroMAG pipeline. ZeroMAG organizes companion modalities around a configuration-invariant adapter, constructs a modality-subject-task condition from unlabeled recordings and task context, and generates adapter weights in a function-constrained latent space learned from source adapters. Across six held-out target datasets and three EFM backbones, ZeroMAG improves balanced accuracy by 7.22 percentage points over EEG-only inference and 4.89 points over direct weight regression, while coming within 0.50 points of supervised multimodal adaptation on average. Ablations further show that removing functional supervision from either representation learning or conditional generation degrades generated-adapter performance, confirming the contribution of both components.