Do Emotion Concepts Generalize Across Sources, Modalities, and Architectures in Vision-Language Models?

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether affective concepts within vision-language models (VLMs) can generalize across data sources, modalities, and architectures. To this end, we construct a multi-source affective stimulus set and extract six Ekman emotion vectors, revealing their shared relational geometry. We further propose an unsupervised cross-architecture alignment method based on universal activations, which combines linear probing with representational space transformations to achieve emotion representation matching. Our findings demonstrate that textual and visual emotion vectors exhibit consistent cross-modal correspondences. Crucially, the aligned consensus vectors generalize in a zero-shot manner to an unseen fourth architecture. This work provides a novel paradigm for the unified understanding of affective representations in VLMs.
📝 Abstract
Recent studies suggest that large language models encode emotion concepts as structured internal representations, but most existing work focuses on text and a single architecture. Therefore, we ask, do emotion concepts generalize across sources, modalities, and architectures in vision--language models (VLMs)? To address this, we construct CMES (Cross-Modal Emotion Stimuli), a multi-source collection of emotion-conditioned stories, real facial expressions, synthetic portraits, and synthetic emotion-evoking scenes. For each stimulus source, we extract a separate set of six Ekman emotion vectors from each of three VLMs. We report four main findings as follows: 1) Image-derived emotion vectors form a low-dimensional geometry similar to that of text-derived vectors. Valence is relatively stable across sources, while arousal varies more. 2) Text- and image-derived emotion vectors have modest cosine similarity but still show held-out cross-modal correspondence. Text-derived vectors can also steer image interpretation. 3) Cross-architecture correspondence remains even when native cosine is near zero. Transformations estimated from generic ImageNet activations recover both correspondence and causal transfer without using the six emotion vectors or their labels. 4) After aligning representations across architectures, we construct a shared emotion subspace that preserves affective geometry and selective steering effects. The corresponding consensus emotion vectors also generalize to a held-out fourth architecture at two model sizes. These results suggest that emotion representations can share relational structure and causal effects across sources, modalities, and architectures, even when individual vector directions differ.
Problem

Research questions and friction points this paper is trying to address.

emotion concepts
vision-language models
cross-modal generalization
cross-architecture
affective representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Cross-Modal Emotion Representations
Cross-Architecture Alignment
Shared Emotion Subspace
Causal Steering
🔎 Similar Papers