BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the absence of a unified, instruction-driven evaluation framework for assessing large language models’ (LLMs) comprehensive capabilities in electroencephalography (EEG) understanding. To this end, we introduce BrainBench—the first standardized benchmark specifically designed for holistic EEG comprehension—encompassing four core tasks, 17 diverse datasets, and tens of thousands of real-world samples. Models are required to generate scientific reports and heterogeneous outputs (numerical, semantic, and human-annotated artifacts) in response to natural language instructions. We propose two analytical paradigms: CodeAct, which leverages autonomous code execution, and BrainAgent, a structured reasoning agent, both integrating multimodal signal processing with LLM-based inference. Through over 100,000 systematic experiments, we evaluate leading models across varying tasks, difficulty levels, and paradigms, revealing nuanced performance differences and establishing a reproducible, multidimensional platform for intelligent EEG analysis.
📝 Abstract
Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.
Problem

Research questions and friction points this paper is trying to address.

EEG understanding
large language models
benchmarking
comprehensive analysis
instruction-conditioned evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

comprehensive EEG understanding
instruction-conditioned benchmark
large language models
multimodal evaluation
structured agentic analysis
🔎 Similar Papers
No similar papers found.