BAT-CLIP: Trimodal Alignment of Brain, Audio and Text

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing brain-speech alignment approaches that anchor to a single modality, which struggle to simultaneously preserve temporal structure and semantic information. To overcome this, we propose the first tri-modal CLIP framework. By leveraging self-supervised foundation models for feature extraction, our method employs contrastive learning to jointly align intracranial electroencephalography (iEEG) neural embeddings with dual anchors in both audio and text modalities. Furthermore, we introduce a shared frozen manifold mechanism to achieve synergistic representations across brain, speech, and text, effectively transcending the constraints of single-anchor alignment. Experiments on a naturalistic podcast benchmark demonstrate that, compared to bi-modal baselines, the proposed framework yields more robust multi-modal representations and significantly enhances decoding performance.
📝 Abstract
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.
Problem

Research questions and friction points this paper is trying to address.

brain-speech alignment
trimodal alignment
CLIP
multimodal speech processing
iEEG decoding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Trimodal Alignment
BAT-CLIP
iEEG
Self-supervised Foundation Models
CLIP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Suhyun Kim
Suhyun Kim
Kyung Hee University
Artificial IntelligenceData ScienceCompilers
J
Jinmo Han
Seoul National University, Seoul, Republic of Korea
D
Danny Dongyeop Han
Seoul National University, Seoul, Republic of Korea
A
Ahhyun Lucy Lee
Seoul National University, Seoul, Republic of Korea
J
Jewoon Lee
Seoul National University, Seoul, Republic of Korea
Y
Yonghyeon Gwon
Seoul National University, Seoul, Republic of Korea
Z
Zach Paris
Dartmouth College, Hanover, NH, USA
C
Chun Kee Chung
Seoul National University, Seoul, Republic of Korea
S
Saewoong Bahk
Seoul National University, Seoul, Republic of Korea
Nam Soo Kim
Nam Soo Kim
Seoul National University, Department of Electrical and Computer Engineering
Seong Jae Hwang
Seong Jae Hwang
Yonsei University
Machine LearningComputer VisionMedical Imaging
Jiook Cha
Jiook Cha
Seoul National University
Human NeuroscienceDevelopmental SciencesMachine Learning