🤖 AI Summary
This study addresses the difficulty of Vision Transformers in effectively grounding abstract concepts without direct visual referents. To this end, it proposes a “metaphorical anchoring” mechanism that leverages concrete, interpretable concepts to bridge visual signals and abstract semantics. Methodologically, Transcoders are applied to parse intermediate features from CLIP and DINO encoders, combined with circuit tracing over an icon dataset. The analysis reveals structured metaphorical circuits wherein perceptual primitives dominate early layers and object-level anchors emerge prior to abstract targets. Furthermore, causal interventions confirm the functional involvement of mediating concepts in abstract concept grounding. Collectively, these findings establish a novel paradigm for understanding abstract reasoning within vision-language models.
📝 Abstract
We study how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. We hypothesize a metonymic grounding mechanism in which abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics. By applying Transcoders on CLIP and DINO vision encoders, we recover intermediate features that can be associated with semantic labels for more concrete concepts, and trace their contributions in circuits underlying abstract concept recognition. Experiments on a carefully curated icon dataset reveal structured metonymic circuits, in which perceptual primitives dominate early layers and object-like anchors precede abstract targets. Images containing rendered text instead recruit a distinct perceptual-to-textual route. Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.