semantic alignment

Designs, implements, and evaluates methods that align semantic content across representational spaces and processing stages—mapping, injecting, or anchoring embeddings and relational structures between modalities (e.g., text, image, graph, or neural embeddings), establishing channel- or decision-level interfaces, and enforcing semantic relational or contrastive losses to preserve meaning while guiding model outputs. This includes building mapping functions (for example fmri-to-embedding or brain-to-semantic mappings), semantic-injection or anchoring modules, alignment-bottleneck channels, and analysis procedures that measure how semantic alignment affects predictions and stability.

semanticalignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

It remains unclear whether multimodal models (e.g., CLIP) better capture experiential semantics and align with human brain fMRI responses compared to unimodal language models. Method: We jointly modeled word representations from both multimodal and large language models against experiential semantic norms and high-resolution fMRI data, conducting cross-modal neural alignment evaluation. Contribution/Results: Contrary to prevailing assumptions, we provide the first empirical evidence that large language models significantly outperform multimodal models in both experiential semantic fidelity and fMRI response prediction accuracy. Their learned representations not only better reflect human experiential cognitive structure but also encode unique semantic dimensions—orthogonal to classical experiential models—yet highly predictive of neural activity. These findings challenge the “multimodality-is-inherently-superior” hypothesis and reveal latent, deep experiential semantic capabilities in language models, offering novel neurocognitive evidence for the cognitive plausibility of linguistic representations.

Assess model alignment with human fMRI responsesCompare multimodal vs language models' experiential information captureEvaluate unique brain-relevant semantic information in models

Brain-aligning of semantic vectors improves neural decoding of visual stimuli

Mar 22, 2024
SV
Shirin Vafaei
🏛️ University of Osaka | Juntendo University | Nara Medical University

Existing neural decoding methods rely on pre-trained image or text feature vectors, yet their semantic structures fundamentally mismatch the brain’s intrinsic neural representations, limiting decoding accuracy. To address this, we propose Brain Alignment—a novel framework that explicitly aligns semantic vector spaces with cortical functional organization via fMRI-supervised fine-tuning. This approach bridges the modality gap by learning brain-grounded semantic embeddings. Crucially, Brain Alignment enables zero-shot transfer to MEG and ECoG modalities, significantly improving stimulus reconstruction correlation (p < 0.001) and fine-grained semantic category classification accuracy across modalities. The gains are robust, generalizable, and reproducible. Furthermore, our analysis reveals that the choice of source semantic space critically determines alignment efficacy—highlighting the importance of architectural compatibility between pretrained features and neural dynamics. This work establishes a principled, brain-informed paradigm for cross-modal neural decoding.

Addresses mismatch between preestablished feature vectors and brain-encoded characteristicsEnhances decoding accuracy across fMRI, MEG, and ECoG neuroimaging datasetsImproves neural decoding by aligning semantic vectors with brain representations

This study investigates the semantic alignment mechanism between vision and language deep models under unsupervised conditions. To this end, we conduct deep representation analysis, cross-modal similarity modeling, Pick-a-Pic forced-choice evaluation, and multi-caption/image matching assessment. Results show that semantic alignment peaks at middle-to-late network layers, exhibiting strong semantic sensitivity and robustness to visual appearance variations. Moreover, averaging representations across multiple instances significantly enhances alignment strength—surpassing conventional one-to-one pairing paradigms and better reflecting human fine-grained preferences in many-to-many image-text scenarios. Key contributions include: (1) the first empirical confirmation that unimodal models encode a shared semantic structure consistent with human judgments; and (2) the discovery that aggregating multiple examples improves alignment quality, with substantial gains achieved while preserving semantic fidelity.

Examining how semantic changes affect cross-modal representational alignmentInvestigating where alignment emerges in vision and language networksTesting whether models capture human preferences in image-text matching

Language models align with brain regions that represent concepts across modalities

Aug 15, 2025
MR
Maria Ryskina
🏛️ Vector Institute for AI | MIT

This study investigates the dissociation between linguistic form representations and conceptual semantic representations in language models, and examines their neural alignment with cross-modal conceptual processing regions in the human brain. Method: We propose a novel metric—“cross-modal semantic consistency”—to quantify the fMRI response consistency across brain regions activated by the same concept presented in three modalities: sentences, word clouds, and images. Leveraging both unimodal language models (LMs) and language-vision multimodal models (LVMs), we systematically evaluate how well their internal representations predict activation patterns in brain regions exhibiting high cross-modal consistency. Results: Both LMs and LVMs significantly outperform language-only task baselines in predicting neural responses in non-linguistically specialized regions—such as the anterior temporal lobe and angular gyrus—demonstrating that these models implicitly encode cross-modal conceptual knowledge. This finding provides novel neuroscientific evidence for the semantic nature of large language models and advances the development of brain–machine semantic interfaces.

Assess if LMs internally represent cross-modal conceptsMeasure meaning consistency across input modalities in brainStudy LM-brain alignment for conceptual meaning representation

The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities

Nov 07, 2024
ZW
Zhaofeng Wu
🏛️ MIT | University of Southern California | Allen Institute for AI

This work investigates whether large language models (LLMs) spontaneously develop a unified, cross-lingual and cross-modal semantic representation space—spanning text, code, images, audio, and arithmetic—and empirically tests the “semantic hub” hypothesis: that intermediate model layers form functionally shared representations analogous to the human brain’s cross-modal semantic hubs. Method: Drawing inspiration from neuroscience’s “hub-and-spoke” model, we integrate logit lens interpretability analysis, cross-modal and cross-lingual embedding similarity metrics, controlled representational interventions, and intermediate-layer probing. Contribution/Results: We find strong semantic alignment across diverse inputs at intermediate layers; moreover, interventions on unimodal representations reliably predict output changes in other modalities—demonstrating that this space is not a training byproduct but actively leveraged for inference. These results uncover intrinsic mechanisms underlying multilingual and multimodal semantic alignment, establishing a novel paradigm for universal intelligent representation modeling.

Interventions in one data type affect others predictably.Model's ability to process diverse data types similarly.Shared semantic representation across languages and modalities.

Latest Papers

What's happening recently
View more

This work addresses the challenges of decoding silent, internal speech—namely the absence of overt output, scarce neural data, and substantial inter-subject variability—by introducing MindAlign, a decoupled two-stage brain-to-language framework. In the first stage, fMRI signals are mapped into a shared multimodal semantic space to produce a semantic sketch. The second stage leverages visual context and a prompting mechanism to guide a frozen multimodal large language model for open-ended text generation, eliminating the need for language model fine-tuning. This approach enables cross-subject decoding without subject-specific adaptation and significantly outperforms both fMRI-only and random baselines. The results demonstrate that neural signals encode semantic information beyond image priors and establish new advances in scalability and generalization for brain-to-text decoding.

brain-to-textfMRIinner speech decoding

This study addresses the challenge of efficiently decoding visual, linguistic, or auditory stimulus representations from fMRI neural activity. To this end, the authors propose a concise yet effective linear contrastive decoding framework that aligns brain activity with the embedding spaces of multimodal foundation models to enable cross-modal mapping. A key finding is that performance gains primarily stem from the contrastive learning objective rather than increased model complexity. Across multiple datasets encompassing images, text, and sounds, the proposed method consistently outperforms ridge regression and nonlinear baselines, demonstrating strong generalization capabilities and validating the efficacy of the alignment paradigm.

brain decodingcontrastive learningfMRI

Current text-to-image generation models struggle to achieve smooth transitions between semantically similar prompts due to substantial differences in token sequences—particularly in wording, ordering, and conceptual positioning—which hinders effective image blending and continuous editing. This work proposes a Token-to-Token Alignment framework that, without modifying the underlying model, employs a two-stage strategy: first aligning the semantic structures of prompts and then aligning their token embedding representations. By reconstructing diverse prompts into a shared structural form, the method reveals that the latent continuous semantic structure within the text embedding space can be effectively leveraged through representation alignment. Consequently, linear interpolation in this aligned space yields coherent semantic transitions, significantly enhancing the quality of image semantic mixing and continuous editing.

embedding spacesemantic blendingsemantic structure

Current evaluations of language models struggle to assess their comprehension of abstract concepts, and the high-dimensional semantic spaces they operate in often lack interpretability. This work introduces topological data analysis into language model evaluation for the first time, proposing a semantic alignment framework that maps low-dimensional, interpretable knowledge structures—such as ontologies and knowledge graphs—onto model embedding spaces. This approach enables cross-lingual and cross-modal tracking of semantic consistency, effectively uncovering the evolutionary dynamics of conceptual representations during model training. Furthermore, it substantially enhances the interpretability of evaluations concerning cross-lingual phrase understanding, offering a principled means to probe how abstract knowledge is encoded and transformed within modern language models.

concept collapsecross-lingual transferembedding spaces

This work addresses the challenge that vision-language models struggle to accurately ground abstract semantics—such as idiomatic meanings of compound nouns—in high-fidelity image generation, where increased visual realism can interfere with compositional semantic understanding. To this end, the authors introduce the DIVA benchmark, which employs diagrammatic images to separately anchor literal and idiomatic interpretations. They further propose, for the first time, architecture-agnostic metrics: a semantic alignment gap (Δ) and a directional bias b(t), to quantify the disparity in visual grounding between these two semantic types. Experiments across eight state-of-the-art models reveal a pervasive literalness bias that persists despite model scaling and intensifies with higher visual fidelity, suggesting that iconographic abstraction enhances symbolic semantic alignment.

compositional understandingidiomatic interpretationsemantic grounding

Hot Scholars

MS

Mubarak Shah

Trustee Chair Professor of Computer Science, University of Central Florida
Computer Vision
XY

Xiaocui Yang

Lecturer, Northeastern University (China)
Multimodal Sentiment AnalysisData MiningMultimodal Large Language Models
JV

Josef van Genabith

DFKI German Research Center for Artificial Intelligence, Saarland University
Natural Language ProcessingMachine TranslationComputational LinguisticsComputational Semantics
HX

Hui Xiong

Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser