🤖 AI Summary
This study addresses the semantic unfaithfulness in text-to-image models caused by metaphorical vehicles being erroneously rendered as visible entities. It formally defines this phenomenon as "metaphorical vehicle intrusion," demonstrating that visual presence does not equate to semantic fidelity. To investigate this issue, we construct VISTA, a multilingual benchmark, and introduce V-Score as a diagnostic metric alongside VISTA-Guard, a mitigation strategy based on lightweight prompt intervention. Furthermore, we propose a role-aware, question-answering-based evaluation method. Extensive experiments confirm that this deficiency is prevalent across languages. Notably, VISTA-Guard significantly reduces the rate of vehicle intrusion, effectively enhancing the metaphorical semantic fidelity of generated images.
📝 Abstract
Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multilingual benchmark of figurative prompts organized by Figurative Form and Mapping Mechanism. We further propose V-Score, a diagnostic question-answering metric that evaluates role-aware figurative faithfulness in generated images. Evaluations on recent high-performing TTI models show that vehicle intrusion persists across languages and figurative categories. As a lightweight mitigation, we introduce VISTA-Guard, which partially reduces vehicle intrusion and suggests a practical path toward more figuratively faithful TTI generation. All resources will be released publicly.