π€ AI Summary
This study investigates whether multimodal foundation models exhibit semantic discrepancies and ambiguity loss between text and image modalities when processing polysemous words. By prompting 17 text-to-image models and 15 text generation models with context-free polysemous terms, the authors systematically analyze the semantic diversity of generated outputs using a human-annotated sense taxonomy and normalized entropy as a metric. The work reveals, for the first time, that image generation substantially compresses the semantic space of polysemous words (normalized entropy: 0.10), markedly lower than both text generation (0.25) and human imagination (0.47). Furthermore, modelsβ self-assessed semantic distributions are significantly richer than those manifested in their actual outputs, exposing a modality-specific gap between internal understanding and external expression.
π Abstract
Human language is highly polysemous. Many common words (e.g., 'bank' or 'palm') carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.