STEAM:ASpatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization and high adaptation costs of existing EEG decoding methods, which struggle to balance universality, accuracy, and efficiency. To overcome these challenges, we propose a dual-branch spatiotemporal encoder framework incorporating a Shared Soft Mixture-of-Experts (SSMoE) mechanism to enable complementary alignment of spatial and temporal representations. Furthermore, we introduce a hierarchical pretraining strategy that facilitates efficient downstream adaptation from a universal initialization. Evaluated across seven datasets and fourteen distinct evaluation settings, our approach achieves state-of-the-art average performance, significantly improving paradigm-specific decoding accuracy while maintaining inference computational cost (FLOPs) within practical limits.
📝 Abstract
Brain-computer interfaces (BCIs) have been widely used in motor rehabilitation, disease diagnosis, and other neural engineering scenarios. However, conventional neural signal decoding algorithms often suffer from limited generalizability and high adaptation costs, motivating recent interest in BCI foundation models. Existing approaches still struggle to jointly achieve general transferability, accurate decoding, and efficient downstream adaptation. We present STEAM, a hierarchical transfer framework that reconciles general-purpose representation learning with paradigm-specific specialization in EEG foundation models. The framework is instantiated as a dual-branch spatio-temporal encoder in which a shared soft mixture-of-experts (SSMoE) module aligns the spatial and temporal branches, allowing complementary representations to exchange information through a compact set of soft slots. Across seven downstream datasets and fourteen evaluation settings, STEAM attains the best average rank among the compared methods at a competitive inference cost measured in FLOPs. Building upon the Stage-I general initialization, the hierarchical pre-training strategy further specializes the model to a target paradigm without retraining from scratch, yielding consistent gains in paradigm-specific decoding accuracy.
Problem

Research questions and friction points this paper is trying to address.

brain-computer interfaces
EEG decoding
transferability
downstream adaptation
generalizability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixture-of-Experts
Spatio-Temporal Alignment
Hierarchical Pre-training
EEG Decoding
Foundation Model
🔎 Similar Papers
💼 Related Jobs
Z
Zhu Chen
Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China; Zhongguancun Academy, Beijing 100094, China
Dingkun Liu
Dingkun Liu
Tsinghua University
brain machine interfaceartificial intelligence
Yuheng Chen
Yuheng Chen
Elmore Family School of Electrical and Computer Engineering, Purdue University
Inverse DesignNanophotonicsMachine LearningSimulation
D
Dongrui Wu
Ministry of Education Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China; Zhongguancun Academy, Beijing 100094, China