NegT2IBench: When Negation Changes the Picture. A Polarity Benchmark for Text-to-Image Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the neglect of negation constraint evaluation in existing text-to-image benchmarks, which hinders assessing models' ability to follow instructions such as "not a red cup." We construct PolarityBench, comprising 4,800 prompts that establish negation constraints as an independent evaluation dimension for the first time, effectively disentangling negation effects from prompt complexity. Furthermore, we design a detector-based automated scoring system enabling fine-grained failure localization with minimal computational overhead. Experiments reveal that most models perform significantly worse on negative than positive instructions, with color attributes exhibiting the most severe failures. Notably, 41.5% of failure cases directly render prohibited content, validating both the effectiveness and diagnostic value of the proposed benchmark.
📝 Abstract
Text-to-image (T2I) models are judged by benchmarks that measure whether requested content appears, but these benchmarks largely overlook the complementary ability to satisfy negated constraints, for example, generating "a non-red cup." Measuring negation raises challenges not faced by affirmation-based benchmarks and requires careful prompt and evaluation design. We introduce NegT2IBench, a benchmark of 4,800 prompts covering two attribute types and four relation categories. Prompts are organized by polarity: the number of positive statements that must hold and negated statements that must not, each ranging from 0 to 2. Varying the two independently separates the effect of negation from the effect of prompt complexity. Our detector-based scoring is reproducible, auditable, and pinpoints which requirement failed. On 600 images with three-annotator labels, it agrees with humans as closely as vision-language judges up to 30x larger, while using only a fraction of their GPU memory. Across eleven T2I models and 211,200 images, nine score lower on a single negated statement than on a single positive one. Per-statement scoring reveals that the loss is largest for color and near zero for proximity, and that 41.5% of failed statements render exactly what the prompt forbids. Rendering what a prompt asks for and withholding what it forbids are distinct capabilities that an aggregate compositional score cannot distinguish. NegT2IBench measures the latter directly, providing a controlled testbed for diagnosing negation failures and developing methods to overcome them.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-Image Benchmark
Negation Understanding
Polarity-based Evaluation
Detector-based Scoring
Compositional Generation
🔎 Similar Papers
No similar papers found.