🤖 AI Summary
This study addresses the limitation of existing brain-speech alignment approaches that anchor to a single modality, which struggle to simultaneously preserve temporal structure and semantic information. To overcome this, we propose the first tri-modal CLIP framework. By leveraging self-supervised foundation models for feature extraction, our method employs contrastive learning to jointly align intracranial electroencephalography (iEEG) neural embeddings with dual anchors in both audio and text modalities. Furthermore, we introduce a shared frozen manifold mechanism to achieve synergistic representations across brain, speech, and text, effectively transcending the constraints of single-anchor alignment. Experiments on a naturalistic podcast benchmark demonstrate that, compared to bi-modal baselines, the proposed framework yields more robust multi-modal representations and significantly enhances decoding performance.
📝 Abstract
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.