Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing text-to-image models, whose OCR-based rewards neglect Chinese character component structures, yielding rendered characters that are visually similar yet structurally inaccurate. We propose IDSpect, a method that leverages Ideographic Description Sequences (IDS) to decompose Chinese characters into component-level units and trains a specialized recognizer for fine-grained visual alignment and reward feedback. Additionally, we design a globally unique token credit mechanism that effectively overcomes the shortcomings of coarse-grained feedback without modifying the generator or increasing inference costs. By post-training Qwen-Image via Group Relative Policy Optimization (GRPO)-based reinforcement learning, our approach achieves state-of-the-art performance on the LongText and GenTextEval benchmarks, demonstrating superior structural quality and semantic alignment in Chinese character rendering.
📝 Abstract
Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic Description Sequences (IDS), which comprise spatial operators and character components, and train an expert IDS recognizer to transcribe rendered Chinese text into this representation. Building on this recognizer, we introduce IDSpect, which deterministically decomposes the target text into IDS tokens and aligns crop-level visual IDS predictions with the target sequence. Globally unique token credit makes this comparison robust to the order of detected text regions. Combined with a whole-character semantic reward, IDSpect supplies fine-grained credit with component and spatial-relation without changing the image generator or adding inference-time cost. Experiments with GRPO post-training of Qwen-Image demonstrate that IDSpect achieves leading structural quality and semantic alignment on LongText and GenTextEval.
Problem

Research questions and friction points this paper is trying to address.

Chinese text rendering
text-to-image models
radical-level errors
fine-grained evaluation
Ideographic Description Sequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ideographic Description Sequences (IDS)
Fine-grained reward
Chinese text rendering
Reinforcement learning
IDSpect
🔎 Similar Papers
No similar papers found.
Y
Yazhen Xie
Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI
X
Xingsong Ye
Institute of Trustworthy Embodied AI, Fudan University; Shanghai Key Laboratory of Multimodal Embodied AI
Zhineng Chen
Zhineng Chen
Institute of Trustworthy Embodied AI, Fudan University
Computer VisionOCRMultimedia Analysis