🤖 AI Summary
This study addresses the underestimation of text-to-image generation performance caused by existing evaluations overlooking high-difficulty compositional prompts. We propose a novel benchmarking paradigm integrating automated complexity filtering with LLM-driven MECE scoring. By establishing an independent evaluation framework to compare four state-of-the-art systems, this approach effectively mitigates biases inherent in traditional metrics. Experimental results demonstrate that Gemini 3 Pro Image achieves the highest score of 84.8, while simultaneously revealing pervasive deficiencies in counting and geometric understanding among leading models. Consequently, this work provides a more precise, reproducible evaluation standard and methodological foundation for assessing advanced text-to-image capabilities under challenging compositional constraints.
📝 Abstract
Text-to-image models are typically reported on average-case prompts, which understates the gap between systems on compositionally demanding requests involving precise object counts, multi-object attribute binding, legible embedded text, and explicit spatial constraints. We evaluate four production text-to-image systems: Hunyuan 3.0, Gemini 3 Pro Image ("Nano Banana Pro"), Black Forest Labs FLUX.2, and Ideogram 3.0. The evaluation uses the 48 hardest prompts drawn from the DataSeeds.AI Sample Dataset (DSD), selected through an automated complexity-scoring pass over the full corpus. Every generated image is graded using an independent-judge rubric. GPT-5.4-Pro authors an atomic, weighted, mutually exclusive and collectively exhaustive (MECE) evaluation rubric, while Gemini 3.1 Pro Preview independently determines whether each criterion is satisfied. Gemini 3 Pro Image ranks first with a score of 84.8/100, narrowly ahead of FLUX.2 at 82.3/100. Ideogram 3.0 and Hunyuan 3.0 score 65.7/100 and 63.3/100, respectively. Failure analysis shows that the leading systems primarily lose points through object miscounting and geometric artifacts, whereas the trailing systems more frequently produce garbled text. Ideogram 3.0 also frequently omits requested elements. Full per-sample rubrics, scores, and failure annotations are available from the authors upon request.