๐ค AI Summary
Text-to-image (T2I) models exhibit weak creativity, poor textโimage alignment, and low persuasiveness when generating advertising images from implicit prompts.
Method: We propose CAPโthe first three-dimensional evaluation framework jointly assessing Creativity, Alignment (prompt fidelity), and Persuasiveness. CAP integrates multi-dimensional human evaluation, implicit-versus-explicit prompt comparison, quantitative measurement of visual-semantic consistency, and behavioral persuasion experiments to systematically uncover structural deficiencies of mainstream T2I models under implicit semantics. We further introduce a lightweight enhancement strategy targeting all three dimensions.
Contribution/Results: CAP provides an interpretable, reproducible, and multi-objective evaluation and optimization paradigm for advertising image generation. Our enhancement strategy yields statistically significant average improvements of 18.7% across all three dimensions (p < 0.01), substantially elevating generation quality.
๐ Abstract
We address the task of advertisement image generation and introduce three evaluation metrics to assess Creativity, prompt Alignment, and Persuasiveness (CAP) in generated advertisement images. Despite recent advancements in Text-to-Image (T2I) generation and their performance in generating high-quality images for explicit descriptions, evaluating these models remains challenging. Existing evaluation methods focus largely on assessing alignment with explicit, detailed descriptions, but evaluating alignment with visually implicit prompts remains an open problem. Additionally, creativity and persuasiveness are essential qualities that enhance the effectiveness of advertisement images, yet are seldom measured. To address this, we propose three novel metrics for evaluating the creativity, alignment, and persuasiveness of generated images. Our findings reveal that current T2I models struggle with creativity, persuasiveness, and alignment when the input text is implicit messages. We further introduce a simple yet effective approach to enhance T2I models' capabilities in producing images that are better aligned, more creative, and more persuasive.