Multimodal LLMs Can Learn to Read Brain Signals: A Vision--Language Model for Unified Multi-Task EEG Decoding

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of generalizing EEG signal decoding across tasks, subjects, and recording sessions. To this end, it proposes BraVista, a novel framework that introduces the first EEG-to-image interface to encode multi-channel EEG signals into structured images. By integrating instruction-conditioned fine-tuning, BraVista leverages the general priors of pre-trained vision-language models (VLMs) to achieve unified multi-task decoding without requiring large-scale EEG-specific pre-training. Experimental evaluations demonstrate that this paradigm achieves superior performance across four distinct tasks, including sleep staging. Furthermore, analyses confirm that the model effectively captures genuine neural information rather than superficial visual artifacts. Overall, this work establishes a promising new pathway for efficient and generalizable EEG decoding by bridging electrophysiological signals with foundational vision-language representations.
📝 Abstract
Learning EEG representations that generalize across cognitive tasks, subjects, and recording conditions remains a key challenge in electroencephalography (EEG) decoding. Recent advances in foundation models have improved EEG decoding performance, yet a fundamental open question remains: how to effectively interface neural signals with these models to enable multi-task learning across datasets. To investigate this question, we introduce BraVista, a visual-language framework that encodes multichannel EEG signals as structured images and enables multi-task learning through instruction-conditioned vision-language models (VLMs). Our approach relies on continued post-training of a general-domain VLM, leveraging its visual and linguistic priors to adapt to neural signals without a separate large-scale EEG-specific pretraining stage. We evaluate BraVista on four datasets spanning sleep staging, emotion recognition, cognitive workload classification, and abnormal EEG detection, showing strong performance across these tasks. Further analyses show that the choice of EEG-to-image representation is critical to performance. Moreover, through controlled perturbations of the EEG signal, we observe a gradual performance degradation under increasing noise, suggesting that the model relies on EEG-relevant information rather than superficial visual patterns. Together, these findings establish structured visual representations as an effective and scalable interface between neural signals and general-domain foundation models for unified multi-task EEG decoding.
Problem

Research questions and friction points this paper is trying to address.

EEG decoding
multi-task learning
foundation models
neural signal interfacing
representation generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

EEG decoding
Vision-language model
Multi-task learning
Structured image representation
Foundation models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Parastoo Azizeddin
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California
O
Omid Sharafi
Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California
Maryam M. Shanechi
Maryam M. Shanechi
Departments of Electrical & Computer Eng., Computer Science, Biomedical Eng., USC
Neural EngineeringMachine LearningBrain-Machine InterfacesControl TheoryNeuroscience