🤖 AI Summary
This work addresses the challenge of evaluating whether purely text-based language models genuinely possess original visual ideation capabilities, as their fluent descriptions may mask non-renderable or clichéd content. The authors introduce the concept of Visual Creative Ideation (VCI) and construct the Ekphrasis benchmark, comprising 400 tasks across four categories—abstraction, composition, transformation, and adaptation—to assess models’ ability to generate textual visual proposals that are useful, expressive, and collectively novel. Innovatively decoupling VCI from linguistic fluency, the study employs a dimension-guided anonymous pairwise comparison framework, Bradley-Terry preference aggregation, and typologized creativity maps to quantify novelty, validating its efficacy through cross-modal rendering. Experiments demonstrate that VCI effectively disentangles three key dimensions, revealing diverse capability profiles among high-scoring models, with text-level VCI rankings showing strong alignment with human image preferences.
📝 Abstract
Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.