Aligning Brain Signals with Multimodal Speech and Vision Embeddings

📅 2025-10-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the neural alignment between brain activity during natural language comprehension and hierarchical representations from multimodal pre-trained models—specifically, auditory (wav2vec 2.0) and vision-language (CLIP) models. Using high-density EEG recorded during auditory sentence processing, we systematically evaluated the predictive power of layer-wise model embeddings for neural responses via ridge regression and contrastive decoding. Our key contribution is a novel “multimodal + layer-aware” representation strategy that integrates cross-model and cross-layer features through concatenation and summation across layers. Results demonstrate that multimodal joint representations significantly outperform unimodal or shallow-layer baselines—particularly in higher-order semantic regions—yielding an average 18.7% increase in explained variance (R²). This provides the first empirical evidence of a neurobiological hierarchy mapping low-level auditory processing onto progressively abstract, cross-modal semantic representations during language understanding.

Technology Category

Natural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Multimodal LearningComputer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
When we hear the word "house", we don't just process sound, we imagine walls, doors, memories. The brain builds meaning through layers, moving from raw acoustics to rich, multimodal associations. Inspired by this, we build on recent work from Meta that aligned EEG signals with averaged wav2vec2 speech embeddings, and ask a deeper question: which layers of pre-trained models best reflect this layered processing in the brain? We compare embeddings from two models: wav2vec2, which encodes sound into language, and CLIP, which maps words to images. Using EEG recorded during natural speech perception, we evaluate how these embeddings align with brain activity using ridge regression and contrastive decoding. We test three strategies: individual layers, progressive concatenation, and progressive summation. The findings suggest that combining multimodal, layer-aware representations may bring us closer to decoding how the brain understands language, not just as sound, but as experience.
Problem

Research questions and friction points this paper is trying to address.

Aligning EEG brain signals with multimodal speech and vision embeddings.
Identifying which pre-trained model layers best reflect brain's layered processing.
Combining multimodal representations to decode language understanding as experience.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Aligning EEG signals with multimodal speech embeddings
Comparing wav2vec2 and CLIP model layer representations
Using ridge regression and contrastive decoding methods
💼 Related Jobs
No related jobs found.
K
Kateryna Shapovalenko
Carnegie Mellon University, Pittsburgh, PA 15213
Q
Quentin Auster
Carnegie Mellon University, Pittsburgh, PA 15213