Score
Designs and implements preprocessing pipelines and mapping models that transform fMRI brain-activation data into fixed-dimensional embedding vectors compatible with external representation spaces. This includes signal denoising and alignment, fitting linear or nonlinear brain-to-embedding mappings, and evaluating embedding alignment and cross-modal retrieval performance across datasets and modalities.
This study systematically investigates the alignment mechanisms between deep neural networks and human brain function, focusing on neural encoding (input → brain response) and decoding (brain signals → semantic/perceptual reconstruction) across language, vision, and audition. Methodologically, it introduces the first cross-modal unified framework integrating fMRI, MEG, and EEG data with Transformers, CNNs, and diffusion models, augmented by canonical correlation analysis (CCA), representational similarity analysis (RSA), and adversarial training; it further proposes a multidimensional evaluation metric for brain-alignment quality and incorporates comparative animal-model evidence and neuroethics considerations. Key contributions include: (1) uncovering strong correspondence between high-level semantic model representations and hierarchical gradients in temporal and prefrontal cortices; (2) achieving high-fidelity image and speech reconstruction, with ImageNet-class decoding accuracy exceeding 85%; and (3) advancing novel paradigms for clinical brain–computer interfaces and early diagnosis of cognitive disorders.
Existing neural decoding methods rely on pre-trained image or text feature vectors, yet their semantic structures fundamentally mismatch the brain’s intrinsic neural representations, limiting decoding accuracy. To address this, we propose Brain Alignment—a novel framework that explicitly aligns semantic vector spaces with cortical functional organization via fMRI-supervised fine-tuning. This approach bridges the modality gap by learning brain-grounded semantic embeddings. Crucially, Brain Alignment enables zero-shot transfer to MEG and ECoG modalities, significantly improving stimulus reconstruction correlation (p < 0.001) and fine-grained semantic category classification accuracy across modalities. The gains are robust, generalizable, and reproducible. Furthermore, our analysis reveals that the choice of source semantic space critically determines alignment efficacy—highlighting the importance of architectural compatibility between pretrained features and neural dynamics. This work establishes a principled, brain-informed paradigm for cross-modal neural decoding.
This study addresses the challenge of improving decoding accuracy and generalizability of visual images from neural signals to advance brain–computer interfaces and deepen understanding of human visual perception. We propose the first human-perception-aligned image encoder, which directly maps multimodal neural signals (e.g., EEG, fMRI) into a semantically consistent visual representation space via cross-modal representation alignment, enabling end-to-end image decoding. Our key contribution is the integration of human behavioral and neural response constraints into the image encoder, substantially enhancing its capacity to model perceptual representations under rapid visual stimulation. Experiments demonstrate up to a 21% improvement in zero-shot image retrieval accuracy. Moreover, the method exhibits robust performance gains across diverse EEG architectures, image encoders, alignment strategies, subjects, and neuroimaging modalities (EEG and fMRI).
This study addresses the challenge of efficiently decoding visual, linguistic, or auditory stimulus representations from fMRI neural activity. To this end, the authors propose a concise yet effective linear contrastive decoding framework that aligns brain activity with the embedding spaces of multimodal foundation models to enable cross-modal mapping. A key finding is that performance gains primarily stem from the contrastive learning objective rather than increased model complexity. Across multiple datasets encompassing images, text, and sounds, the proposed method consistently outperforms ridge regression and nonlinear baselines, demonstrating strong generalization capabilities and validating the efficacy of the alignment paradigm.
This study investigates the neural alignment between brain activity during natural language comprehension and hierarchical representations from multimodal pre-trained models—specifically, auditory (wav2vec 2.0) and vision-language (CLIP) models. Using high-density EEG recorded during auditory sentence processing, we systematically evaluated the predictive power of layer-wise model embeddings for neural responses via ridge regression and contrastive decoding. Our key contribution is a novel “multimodal + layer-aware” representation strategy that integrates cross-model and cross-layer features through concatenation and summation across layers. Results demonstrate that multimodal joint representations significantly outperform unimodal or shallow-layer baselines—particularly in higher-order semantic regions—yielding an average 18.7% increase in explained variance (R²). This provides the first empirical evidence of a neurobiological hierarchy mapping low-level auditory processing onto progressively abstract, cross-modal semantic representations during language understanding.
Low reproducibility of fMRI statistical maps stems from variability across preprocessing and analysis pipelines. To address this, we propose the first unsupervised multi-domain diffusion framework for 3D neuroimaging statistical maps, modeling analytical pipelines as transferable “styles.” Our method integrates auxiliary classifier-guided latent-space constraints with an improved diffusion sampling strategy to enable high-fidelity cross-pipeline style transfer. Compared to GAN-based baselines, our approach preserves anatomical integrity of brain activation patterns while significantly reducing analysis-induced variability (p < 0.01) and enhancing inter-site data consistency. The framework provides a novel tool for fMRI data augmentation, methodological standardization, and clinically interpretable analysis—advancing robustness and reproducibility in neuroimaging research.
This work addresses the challenge of low sample efficiency in fMRI-based encoding and decoding models, which stems from the scarcity of paired fMRI–stimulus data and substantial inter-subject variability. To overcome this, the authors propose a lightweight latent embedding alignment framework that operates with frozen pre-trained encoders and decoders. By leveraging abundant unpaired stimulus embeddings through a reverse semi-supervised learning strategy, the method introduces a novel meta-transfer mechanism that integrates residual debiasing and sparse aggregation to enable effective cross-subject knowledge transfer and alignment refinement. Theoretical analysis provides generalization bounds and safety guarantees under limited sample regimes. Extensive experiments on large-scale fMRI image reconstruction benchmarks demonstrate significant improvements in both sample efficiency and reconstruction performance, confirming the approach’s effectiveness and robustness.
This study addresses the neurodecoding challenge of directly reconstructing co-speech gestures from fMRI signals—a task hindered by the absence of paired {brain signal–speech–gesture} data. To overcome this, we propose a dual-path brain decoding alignment framework that leverages text as a semantic bridge: it jointly optimizes an fMRI-to-text decoder and a text-to-gesture generator, enabling self-supervised, unpaired cross-modal mapping from fMRI to gesture. Multimodal fusion is achieved via region-of-interest (ROI)-based feature analysis and self-supervised alignment. Experimentally, we achieve the first successful generation of expressive, temporally aligned co-speech gestures from natural speech–evoked fMRI data. Furthermore, our analyses reveal distinct functional contributions of motor, language, and default mode networks to gesture generation. This work establishes a novel paradigm for both brain–computer interfaces and the investigation of speech–gesture coupling in cognitive neuroscience.
This work addresses the limitation of existing visual decoding approaches, which predominantly focus on high-level semantics while neglecting pixel-level details, thereby failing to fully capture the brain’s encoding mechanisms of visual information. To overcome this, the authors propose a hierarchical alignment strategy that integrates multi-scale pre-trained visual encoders, coupled with a contrastive learning objective and a newly designed Fusion Prior mechanism. This approach effectively enhances cross-modal distributional consistency between neural signals and image representations. The method achieves state-of-the-art performance in both quantitative and qualitative evaluations, significantly improving reconstruction fidelity without compromising retrieval accuracy. Notably, it represents the first effort to successfully balance semantic correctness with fine-grained detail preservation in brain-to-image reconstruction.
Analyzing subtle, spatially sparse, and viewpoint-dependent lesions in brain MRI remains challenging. Method: We introduce the largest publicly available paired MRI–clinical report dataset to date (80K samples, a 10× scale-up), and propose a multi-view alignment representation learning framework featuring: (i) a novel implicit query-feature matching mechanism; (ii) quality- and diversity-driven multi-view embedding alignment; and (iii) integration of 3D slice-level feature disentanglement/aggregation with a document-retrieval-inspired cross-modal pretraining paradigm. Contribution/Results: Our approach achieves state-of-the-art performance across both vision-language understanding and pure-vision medical tasks. We open-source the BRAT foundation model, enabling zero-shot transfer and clinical report generation.
This study addresses the challenge of limited training data in electrocorticography (ECoG)-based neural encoding models, which stems from the scarcity of implantable patients. To overcome this limitation, the authors propose fine-tuning language representation models using non-invasive functional magnetic resonance imaging (fMRI) data and transferring the learned representations to model high spatiotemporal resolution ECoG signals. The work demonstrates for the first time that fMRI data—despite its temporal resolution being two orders of magnitude lower than ECoG—can significantly enhance ECoG prediction performance, thereby validating the efficacy of cross-modal transfer and data augmentation. Experimental results show that the fine-tuned models yield substantially improved ECoG predictions, with performance consistently increasing as more fMRI data are incorporated, and maintain robust generalization even under temporally downsampled conditions.