🤖 AI Summary
This study addresses the limited transferability of EEG foundation models caused by shortcut learning from positional cues. To mitigate this issue, we propose the NSP framework, which constrains prediction targets and contexts to suppress reliance on low-information pathways. Furthermore, the framework introduces EMA-based latent supervision, identity residualization, and a topology-disentangled context mechanism to enable efficient masked pretraining. Experimental results demonstrate that the proposed method achieves a macro-balanced accuracy of 63.94% across 30 downstream tasks, outperforming the baseline by 2.35 percentage points. These findings indicate that our approach significantly enhances the transferability of learned EEG representations, offering a robust solution for scalable brain signal modeling.
📝 Abstract
EEG foundation models increasingly use masked prediction to learn from unlabeled recordings, but optimizing this objective does not ensure transferable neural representations. A central challenge is that stable positional cues and local correlations can make masked regions predictable without integrating distributed neural context. To reduce this reliance on low-information prediction paths, we introduce Neural State Prediction (NSP), a latent-predictive framework that constrains both the prediction target and the available context. NSP uses a Target Encoder updated by an exponential moving average (EMA) to define latent supervision. Identity residualization removes additive effects associated with channel identity and relative time from the targets, while topology-separated context excludes their immediate spatial and temporal neighborhood from the visible input. We pretrain NSP on 2.2 million EEG segments from TUEG and evaluate it across 30 downstream datasets spanning clinical diagnosis, sleep staging, emotion recognition, motor imagery, event-related potentials, cognitive-state decoding, and language retrieval. Under full-parameter multi-task fine-tuning on EEG-FM-Bench, NSP achieves 63.94 macro balanced accuracy across 14 datasets, exceeding the strongest evaluated baseline by 2.35 percentage points. Controlled component ablations assess the contribution of each mechanism, while matched context controls and held-out interventions characterize the role of context geometry, signal content, and positional information. Jointly designing latent targets and their context offers a promising direction for EEG foundation models that learn from distributed signal structure.