Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that visual discriminability exceeds linguistic addressability, hindering models from precisely accessing subtle geometric differences in frozen image representations via text. To tackle this, we introduce the FactorAtlas testbed to decouple these two capabilities and derive an image-side contrastive mechanism that narrows the access gap. Furthermore, we propose a matching-based visual grounding method demonstrating that alignment between the two capacities is unnecessary, supported by direction-specificity and ablation control experiments to verify the source of gains. By integrating cross-testing, text-to-image retrieval, and compositional retrieval techniques, this work effectively bridges the inherent textual access gap. The resulting improvements remain consistently stable across compositional retrieval and natural image scenarios, confirming the probing role of visual grounding.
📝 Abstract
Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.
Problem

Research questions and friction points this paper is trying to address.

Vision-language models
Visual discriminability
Linguistic addressability
Text-to-image retrieval
Access gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual discriminability
linguistic addressability
FactorAtlas
matched visual grounding
vision-language models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Woosang Jeon
Seoul National University
J
Jiwon Yang
Seoul National University
S
Soo Chung
Seoul National University
Taehyeong Kim
Taehyeong Kim
Seoul National University
Artificial IntelligenceCognitive ScienceIntelligent Agriculture