π€ AI Summary
This study addresses the evaluation challenges of rendering dense text in image generation across long strings, multiple regions, and complex scenes. To overcome the limitations of existing short-text benchmarks, this work proposes the first Chinese-English bilingual benchmark comprising 432 real-world prompts organized into a three-tier difficulty hierarchy. Methodologically, it introduces Q-Judger, a vision-language model for automated assessment, complemented by human verification to quantify performance across multiple dimensions such as fidelity and legibility. Experimental results reveal substantial performance degradation in mainstream models under high-difficulty conditions; for instance, Qwenβs overall English score drops from 86.5 to 42.86. Ultimately, this research provides a critical evaluation tool and novel insights for advancing dense visual text generation.
π Abstract
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.