What masking geometry works best for EEG foundation models?

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of independent ablations and unclear optimal configurations for masking strategies in EEG foundation model pretraining. Within a unified pipeline, we conduct a systematic ablation of spatiotemporal masking strategies under MAE and JEPA frameworks. By isolating the effects of mask geometry for the first time, we evaluate the linear probing performance of 58 models across 12 datasets. Our analysis reveals an optimal masking configuration shared by both frameworks and identifies a bias inflation collapse mechanism unique to JEPA. The resulting robust masking configuration achieves downstream performance comparable to REVE at substantially lower computational cost, providing critical design guidelines for self-supervised EEG pretraining.
📝 Abstract
EEG foundation models hold promise for scalable brain-signal decoding across clinical and cognitive neuroscience applications, yet their pre-training pipelines remain poorly understood. Among design choices, the masking strategy is particularly critical: it determines what the network must predict and from which context. Yet it has never been ablated in isolation, as each new model bundles a new masking strategy with a new backbone and objective. In this paper, we formalize the design choices for spatio-temporal masking strategies and train various models with a single pipeline under varying masking configurations across two SSL frameworks (MAE and JEPA). We then systematically evaluate the resulting 58 pre-trained models on the 12 datasets of OpenEEGBench under a linear probe. Both frameworks agree on an optimal masking configuration and on shared failure modes. Outside these, performance is robust: 11 MAE and 9 JEPA configurations are statistically indistinguishable from the best. We further identify a novel JEPA-specific failure mode, tagged bias-inflation collapse, invisible to standard detectors. With a well-chosen mask, our pipeline reaches REVE-level downstream performance at a fraction of REVE's pre-training compute.
Problem

Research questions and friction points this paper is trying to address.

EEG foundation models
masking strategy
self-supervised learning
pre-training
spatio-temporal masking
Innovation

Methods, ideas, or system contributions that make the work stand out.

EEG foundation models
spatio-temporal masking
self-supervised learning
bias-inflation collapse
JEPA
🔎 Similar Papers
No similar papers found.