🤖 AI Summary
This study addresses the challenges of designing EEG decoding architectures for multi-task scenarios, which typically incur prohibitive training costs. To this end, we propose an LLM-driven dual-agent neural architecture search framework. This framework introduces a pool-guided architecture discovery mechanism and a performance estimation based on early learning curves (PEEK) strategy, synergistically optimizing architecture generation and early screening to substantially reduce full-training overhead. Experimental results demonstrate that the proposed method achieves an average balanced accuracy of 64.16% across 14 EEG datasets while reducing performance estimation error by 38.1%. These findings confirm the efficacy of the framework in enabling efficient and automated search for EEG decoding models.
📝 Abstract
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection of EEG decoding architectures across tasks. The Designer Agent performs Pool-Guided Architecture Discovery (PGAD), generating and refining architectures through training and validation across multiple EEG tasks, such as emotion recognition, motor imagery, and sleep staging. The Forecaster Agent performs Performance Estimation from Early Knowledge (PEEK), using architecture code, the training protocol, and early learning curves to predict full-budget validation performance and select promising candidates for continued training. Across 14 EEG datasets spanning motor imagery, emotion recognition, and sleep staging, we evaluate AutoBCI with six LLMs, including Opus 5.5 and GPT 5.6 Sol, and compare the architectures selected by the search procedure against ten baselines: six conventional EEG models and four foundation models. The architecture discovered by AutoBCI with Claude Opus 5.5 achieves 64.16% average test balanced accuracy (bAcc), compared with 63.87% for REVE, the strongest baseline on this metric. Using ten observed epochs, PEEK reduces mean absolute error in predicting average validation bAcc from 2.20 to 1.36 percentage points, a 38.1% reduction relative to the best-observed-score baseline.