Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the unclear advantages and lack of systematic evaluation of EEG foundation models by introducing EEG-Arena, a large-scale benchmarking platform. Employing a multi-protocol framework, it systematically compares 30 foundation models against 25 supervised baselines across 57 tasks. The research provides the first empirical evidence that scaling data volume yields greater performance improvements than scaling model parameters, identifying pretraining data expansion as the primary direction for future development. Furthermore, the results demonstrate that EEG foundation models significantly outperform conventional supervised methods. To facilitate reproducible research, the proposed platform is made publicly available as open source.
πŸ“ Abstract
Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG FMs and 25 supervised baselines evaluated on 57 downstream tasks from 23 public datasets. Through more than 20,000 evaluations across five experimental protocols, we assess downstream performance, pretraining benefits, model size scaling, pretraining data scaling, and robustness to channel configuration. We find that (1) EEG FMs outperform strong task-specific supervised baselines on most evaluated tasks, particularly under non-bipolar settings; (2) compared with architecture-matched supervised training from scratch, pretraining improves both early optimization and final downstream performance, with larger and more consistent gains as more labeled downstream data become available; (3) existing EEG FMs do not exhibit a consistent positive relationship between parameter count and downstream performance; (4) under a fixed architecture, increasing the pretraining data scale yields sustained downstream gains; and (5) channel-flexible FMs achieve higher absolute performance than channel-constrained models across most evaluated channel configurations. Together, these findings demonstrate the downstream value of EEG FMs and identify pretraining data expansion as a promising direction for further progress. To support continued research, we release EEG-Arena as an open-source evaluation framework that provides shared infrastructure for reproducible benchmarking, model comparison, and community-driven development.
Problem

Research questions and friction points this paper is trying to address.

EEG foundation models
benchmarking
supervised baselines
scaling
downstream performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

EEG foundation models
benchmarking framework
pretraining data scaling
channel flexibility
open-source evaluation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Z
Zhige Chen
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University
S
Shu Peng
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University
Chengxuan Qin
Chengxuan Qin
University of Liverpool
Machine LearningReinforcement Learning
R
Rui Liu
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University
Rui Yang
Rui Yang
Xi’an Jiaotong-Liverpool University
Data driven fault diagnosistransfer learningdomain adaptationEEG signal classication
K
Kay Chen Tan
Department of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University
Jibin Wu
Jibin Wu
The Hong Kong Polytechnic University
Spiking Neural NetworkNeuromorphic ComputingSpeech ProcessingCognitive Modelling