π€ AI Summary
This study addresses the unclear advantages and lack of systematic evaluation of EEG foundation models by introducing EEG-Arena, a large-scale benchmarking platform. Employing a multi-protocol framework, it systematically compares 30 foundation models against 25 supervised baselines across 57 tasks. The research provides the first empirical evidence that scaling data volume yields greater performance improvements than scaling model parameters, identifying pretraining data expansion as the primary direction for future development. Furthermore, the results demonstrate that EEG foundation models significantly outperform conventional supervised methods. To facilitate reproducible research, the proposed platform is made publicly available as open source.
π Abstract
Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG FMs and 25 supervised baselines evaluated on 57 downstream tasks from 23 public datasets. Through more than 20,000 evaluations across five experimental protocols, we assess downstream performance, pretraining benefits, model size scaling, pretraining data scaling, and robustness to channel configuration. We find that (1) EEG FMs outperform strong task-specific supervised baselines on most evaluated tasks, particularly under non-bipolar settings; (2) compared with architecture-matched supervised training from scratch, pretraining improves both early optimization and final downstream performance, with larger and more consistent gains as more labeled downstream data become available; (3) existing EEG FMs do not exhibit a consistent positive relationship between parameter count and downstream performance; (4) under a fixed architecture, increasing the pretraining data scale yields sustained downstream gains; and (5) channel-flexible FMs achieve higher absolute performance than channel-constrained models across most evaluated channel configurations. Together, these findings demonstrate the downstream value of EEG FMs and identify pretraining data expansion as a promising direction for further progress. To support continued research, we release EEG-Arena as an open-source evaluation framework that provides shared infrastructure for reproducible benchmarking, model comparison, and community-driven development.