Score
Designs, implements, and evaluates methods that align semantic content across representational spaces and processing stages—mapping, injecting, or anchoring embeddings and relational structures between modalities (e.g., text, image, graph, or neural embeddings), establishing channel- or decision-level interfaces, and enforcing semantic relational or contrastive losses to preserve meaning while guiding model outputs. This includes building mapping functions (for example fmri-to-embedding or brain-to-semantic mappings), semantic-injection or anchoring modules, alignment-bottleneck channels, and analysis procedures that measure how semantic alignment affects predictions and stability.
This study investigates the representational potential of foundation models for cross-modal alignment—specifically, whether their unimodal representations inherently capture task-specific semantics and exhibit cross-modal transferability. Method: We formalize “representational potential” and systematically analyze structural regularities and semantic consistency across vision, language, and speech foundation models. Leveraging cross-modal similarity metrics, representation visualization, and neuroscience-inspired evaluation protocols, we assess generalizability and unification capacity across diverse architectures. Contribution/Results: Empirical results demonstrate that pretrained foundation models implicitly acquire semantic invariances requisite for cross-modal alignment—even when trained on unimodal data—thereby exhibiting strong potential as unified multimodal representation backbones. Our work establishes a theoretical framework for cross-modal alignment and introduces a reproducible, multi-faceted evaluation paradigm grounded in representational analysis.
It remains unclear whether multimodal models (e.g., CLIP) better capture experiential semantics and align with human brain fMRI responses compared to unimodal language models. Method: We jointly modeled word representations from both multimodal and large language models against experiential semantic norms and high-resolution fMRI data, conducting cross-modal neural alignment evaluation. Contribution/Results: Contrary to prevailing assumptions, we provide the first empirical evidence that large language models significantly outperform multimodal models in both experiential semantic fidelity and fMRI response prediction accuracy. Their learned representations not only better reflect human experiential cognitive structure but also encode unique semantic dimensions—orthogonal to classical experiential models—yet highly predictive of neural activity. These findings challenge the “multimodality-is-inherently-superior” hypothesis and reveal latent, deep experiential semantic capabilities in language models, offering novel neurocognitive evidence for the cognitive plausibility of linguistic representations.
Existing neural decoding methods rely on pre-trained image or text feature vectors, yet their semantic structures fundamentally mismatch the brain’s intrinsic neural representations, limiting decoding accuracy. To address this, we propose Brain Alignment—a novel framework that explicitly aligns semantic vector spaces with cortical functional organization via fMRI-supervised fine-tuning. This approach bridges the modality gap by learning brain-grounded semantic embeddings. Crucially, Brain Alignment enables zero-shot transfer to MEG and ECoG modalities, significantly improving stimulus reconstruction correlation (p < 0.001) and fine-grained semantic category classification accuracy across modalities. The gains are robust, generalizable, and reproducible. Furthermore, our analysis reveals that the choice of source semantic space critically determines alignment efficacy—highlighting the importance of architectural compatibility between pretrained features and neural dynamics. This work establishes a principled, brain-informed paradigm for cross-modal neural decoding.
This study investigates the semantic alignment mechanism between vision and language deep models under unsupervised conditions. To this end, we conduct deep representation analysis, cross-modal similarity modeling, Pick-a-Pic forced-choice evaluation, and multi-caption/image matching assessment. Results show that semantic alignment peaks at middle-to-late network layers, exhibiting strong semantic sensitivity and robustness to visual appearance variations. Moreover, averaging representations across multiple instances significantly enhances alignment strength—surpassing conventional one-to-one pairing paradigms and better reflecting human fine-grained preferences in many-to-many image-text scenarios. Key contributions include: (1) the first empirical confirmation that unimodal models encode a shared semantic structure consistent with human judgments; and (2) the discovery that aggregating multiple examples improves alignment quality, with substantial gains achieved while preserving semantic fidelity.
This study investigates the dissociation between linguistic form representations and conceptual semantic representations in language models, and examines their neural alignment with cross-modal conceptual processing regions in the human brain. Method: We propose a novel metric—“cross-modal semantic consistency”—to quantify the fMRI response consistency across brain regions activated by the same concept presented in three modalities: sentences, word clouds, and images. Leveraging both unimodal language models (LMs) and language-vision multimodal models (LVMs), we systematically evaluate how well their internal representations predict activation patterns in brain regions exhibiting high cross-modal consistency. Results: Both LMs and LVMs significantly outperform language-only task baselines in predicting neural responses in non-linguistically specialized regions—such as the anterior temporal lobe and angular gyrus—demonstrating that these models implicitly encode cross-modal conceptual knowledge. This finding provides novel neuroscientific evidence for the semantic nature of large language models and advances the development of brain–machine semantic interfaces.
This work investigates whether large language models (LLMs) spontaneously develop a unified, cross-lingual and cross-modal semantic representation space—spanning text, code, images, audio, and arithmetic—and empirically tests the “semantic hub” hypothesis: that intermediate model layers form functionally shared representations analogous to the human brain’s cross-modal semantic hubs. Method: Drawing inspiration from neuroscience’s “hub-and-spoke” model, we integrate logit lens interpretability analysis, cross-modal and cross-lingual embedding similarity metrics, controlled representational interventions, and intermediate-layer probing. Contribution/Results: We find strong semantic alignment across diverse inputs at intermediate layers; moreover, interventions on unimodal representations reliably predict output changes in other modalities—demonstrating that this space is not a training byproduct but actively leveraged for inference. These results uncover intrinsic mechanisms underlying multilingual and multimodal semantic alignment, establishing a novel paradigm for universal intelligent representation modeling.
This work addresses the challenges of decoding silent, internal speech—namely the absence of overt output, scarce neural data, and substantial inter-subject variability—by introducing MindAlign, a decoupled two-stage brain-to-language framework. In the first stage, fMRI signals are mapped into a shared multimodal semantic space to produce a semantic sketch. The second stage leverages visual context and a prompting mechanism to guide a frozen multimodal large language model for open-ended text generation, eliminating the need for language model fine-tuning. This approach enables cross-subject decoding without subject-specific adaptation and significantly outperforms both fMRI-only and random baselines. The results demonstrate that neural signals encode semantic information beyond image priors and establish new advances in scalability and generalization for brain-to-text decoding.
This study addresses the challenge of efficiently decoding visual, linguistic, or auditory stimulus representations from fMRI neural activity. To this end, the authors propose a concise yet effective linear contrastive decoding framework that aligns brain activity with the embedding spaces of multimodal foundation models to enable cross-modal mapping. A key finding is that performance gains primarily stem from the contrastive learning objective rather than increased model complexity. Across multiple datasets encompassing images, text, and sounds, the proposed method consistently outperforms ridge regression and nonlinear baselines, demonstrating strong generalization capabilities and validating the efficacy of the alignment paradigm.
Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.
Current evaluations of language models struggle to assess their comprehension of abstract concepts, and the high-dimensional semantic spaces they operate in often lack interpretability. This work introduces topological data analysis into language model evaluation for the first time, proposing a semantic alignment framework that maps low-dimensional, interpretable knowledge structures—such as ontologies and knowledge graphs—onto model embedding spaces. This approach enables cross-lingual and cross-modal tracking of semantic consistency, effectively uncovering the evolutionary dynamics of conceptual representations during model training. Furthermore, it substantially enhances the interpretability of evaluations concerning cross-lingual phrase understanding, offering a principled means to probe how abstract knowledge is encoded and transformed within modern language models.
This work addresses the challenge that vision-language models struggle to accurately ground abstract semantics—such as idiomatic meanings of compound nouns—in high-fidelity image generation, where increased visual realism can interfere with compositional semantic understanding. To this end, the authors introduce the DIVA benchmark, which employs diagrammatic images to separately anchor literal and idiomatic interpretations. They further propose, for the first time, architecture-agnostic metrics: a semantic alignment gap (Δ) and a directional bias b(t), to quantify the disparity in visual grounding between these two semantic types. Experiments across eight state-of-the-art models reveal a pervasive literalness bias that persists despite model scaling and intensifies with higher visual fidelity, suggesting that iconographic abstraction enhances symbolic semantic alignment.