🤖 AI Summary
This study addresses the limitation of existing text-to-image models, whose OCR-based rewards neglect Chinese character component structures, yielding rendered characters that are visually similar yet structurally inaccurate. We propose IDSpect, a method that leverages Ideographic Description Sequences (IDS) to decompose Chinese characters into component-level units and trains a specialized recognizer for fine-grained visual alignment and reward feedback. Additionally, we design a globally unique token credit mechanism that effectively overcomes the shortcomings of coarse-grained feedback without modifying the generator or increasing inference costs. By post-training Qwen-Image via Group Relative Policy Optimization (GRPO)-based reinforcement learning, our approach achieves state-of-the-art performance on the LongText and GenTextEval benchmarks, demonstrating superior structural quality and semantic alignment in Chinese character rendering.
📝 Abstract
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.