UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

πŸ“… 2026-10-07
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the evaluation challenges of rendering dense text in image generation across long strings, multiple regions, and complex scenes. To overcome the limitations of existing short-text benchmarks, this work proposes the first Chinese-English bilingual benchmark comprising 432 real-world prompts organized into a three-tier difficulty hierarchy. Methodologically, it introduces Q-Judger, a vision-language model for automated assessment, complemented by human verification to quantify performance across multiple dimensions such as fidelity and legibility. Experimental results reveal substantial performance degradation in mainstream models under high-difficulty conditions; for instance, Qwen’s overall English score drops from 86.5 to 42.86. Ultimately, this research provides a critical evaluation tool and novel insights for advancing dense visual text generation.
πŸ“ Abstract
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: https://github.com/LINs-lab/UltraText_Bench.
Problem

Research questions and friction points this paper is trying to address.

visual text rendering
image generation
benchmark
dense text
bilingual evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dense Visual Text Rendering
Bilingual Benchmark
Automated Evaluation
Vision-Language Model
Image Generation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
D
Deyuan Liu
Westlake University
Yihao Hu
Yihao Hu
Duke University
DatabaseQuery OptimizationData Discovery
J
Jingxuan Zhang
Westlake University
X
Xingying Li
Westlake University
Jun Xie
Jun Xie
Southwest Jiao Tong University
Transportation
Jiacheng Liu
Jiacheng Liu
MBZUAI
Machine Learning
J
Jungang Li
HKUST
Y
Yu Huang
CityU
X
Xuanyi Liu
Peking University
Y
Yue Ding
CASIA
Z
Zecheng Wang
Wechat AI
L
Lei Zhao
Westlake University
M
Mingda Wang
Westlake University
Zhenglin Cheng
Zhenglin Cheng
Zhejiang University & Westlake University, SII
Multimodal LearningDiffusion Models
P
Peng Sun
Westlake University, Zhejiang University, Shanghai Innovation Institute
T
Tao Lin
Westlake University