Many Brains, One Geometry: A Shared Visual-Semantic Space for Cross-Dataset fMRI Decoding

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the data silo problem arising from cross-subject and cross-experiment variability in fMRI-based visual decoding by proposing the BRAID-fMRI framework. This framework introduces a novel shared CLIP supervision paradigm that employs an ROI-wise Transformer to align heterogeneous brain activity into a unified semantic space. Furthermore, it leverages multi-positive contrastive learning to handle repeated stimuli, thereby facilitating subject-agnostic modeling and cross-dataset generalization. Experimental results demonstrate that the model achieves a Top-10 accuracy of 35.0%, while expanding the source data pool improves target-set performance by 92.1%. Additionally, the framework effectively preserves hierarchical semantic structures and selectively focuses on the ventral visual cortex, highlighting its capacity for robust and biologically plausible visual decoding across diverse neuroimaging datasets.
πŸ“ Abstract
Visual decoding from fMRI is typically siloed by participant and experiment, obscuring whether heterogeneous neural measurements can be organized within a common computational geometry. Here we introduce BRAID-fMRI (Brain Representation Alignment across Individuals and Datasets), a shared CLIP-supervised decoding framework. BRAID-fMRI uses a single ROI-wise Transformer with optional participant conditioning across eight visual-fMRI datasets comprising 93 dataset-specific participant entries, 430,007 single-trial responses and 162,839 unique stimuli. Regional brain activity is aligned with 512-dimensional CLIP ViT-B/32 representations using a multi-positive contrastive objective that treats repeated stimuli across participants and datasets as positives. BRAID-fMRI supports retrieval across seven evaluation datasets. On eight matched participant entries, it achieves 35.0 +/- 11.1% Top-10 accuracy, exceeding the observed mean accuracy of the two evaluated baselines - the MindEye-style pooled-CLIP decoder (27.1 +/- 5.1%) and ridge regression (21.1 +/- 9.1%) - and attaining the highest observed accuracy for seven of eight entries. In separately trained participant-agnostic models, expanding the source pool increased target-dataset holdout accuracy by up to 92.1% relative to the initial source-training condition. The learned space preserves graded semantic structure, while ablations and saliency highlight ventral and early visual cortex and category-specific motion and attentional systems. These results support scalable cross-dataset decoding into a common CLIP-aligned space, with model sensitivity concentrated in ventral and early visual inputs.
Problem

Research questions and friction points this paper is trying to address.

fMRI decoding
cross-dataset
visual-semantic space
neural representation alignment
cross-participant
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-dataset fMRI decoding
Shared visual-semantic space
Multi-positive contrastive learning
ROI-wise Transformer
CLIP alignment