Simile Understanding in Text-to-Image Models: An Evaluation Framework

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of text-to-image models to misinterpret simile prompts by treating figurative vehicles as literal objects, revealing a semantic gap between rhetorical comprehension and visual grounding. The study introduces the first systematic evaluation framework for simile understanding in generative models, combining template-based controllable data generation, YOLO-based automatic metrics for visual grounding, and Diffusion Lens analysis of intermediate representations in text encoders. Through comprehensive experiments on mainstream diffusion models, the research uncovers pervasive literal-generation errors and establishes a scalable evaluation paradigm that offers concrete directions for improving the visual representation of figurative language.
📝 Abstract
Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object. These systematic failures reveal a gap between figurative language and object-level visual grounding in t2i models. To investigate this issue, we propose a scalable evaluation framework for simile understanding. Our framework includes (1) a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates, (2) automatic grounding metrics based on YOLO (You Only Look Once) detection, and (3) text encoder layer analysis using Diffusion Lens to track how metaphorical vehicles emerge during generation. Experiments across architecturally diverse t2i models reveal consistent literalization failure patterns. We further discuss potential mitigation strategies for improving simile grounding in t2i models.
Problem

Research questions and friction points this paper is trying to address.

simile understanding
text-to-image models
figurative language
visual grounding
metaphorical vehicle
Innovation

Methods, ideas, or system contributions that make the work stand out.

simile understanding
text-to-image models
visual grounding
evaluation framework
figurative language