🤖 AI Summary
This study addresses the challenge that visual discriminability exceeds linguistic addressability, hindering models from precisely accessing subtle geometric differences in frozen image representations via text. To tackle this, we introduce the FactorAtlas testbed to decouple these two capabilities and derive an image-side contrastive mechanism that narrows the access gap. Furthermore, we propose a matching-based visual grounding method demonstrating that alignment between the two capacities is unnecessary, supported by direction-specificity and ablation control experiments to verify the source of gains. By integrating cross-testing, text-to-image retrieval, and compositional retrieval techniques, this work effectively bridges the inherent textual access gap. The resulting improvements remain consistently stable across compositional retrieval and natural image scenarios, confirming the probing role of visual grounding.
📝 Abstract
Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.