🤖 AI Summary
Text-to-image diffusion models achieve high visual fidelity but struggle to faithfully render fine-grained subject details—such as precise text spelling—due to the limited semantic expressivity of standard text tokenizers. To address this, we propose a reference-guided generation framework based on lightweight expert plug-ins: leveraging a reference image as visual conditioning, jointly with textual input, to steer the Stable Diffusion architecture and overcome the representational bottlenecks of language-only conditioning. Our method integrates multiple task-specific, compact expert plug-ins (each with only 28.55M parameters), an auxiliary network, and dedicated loss functions tailored for cross-lingual text rendering and domain-specific applications like logo generation. Extensive experiments demonstrate consistent and significant improvements over existing state-of-the-art methods across English and multilingual text-to-image synthesis, as well as logo generation benchmarks.
📝 Abstract
Text-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating high-fidelity visuals, text-to-image models can still struggle with precisely rendering subjects, such as text spelling. To address this challenge, this paper explores using additional conditions of an image that provides visual guidance of the particular subjects for diffusion models to generate. In addition, this reference condition empowers the model to be conditioned in ways that the vocabularies of the text tokenizer cannot adequately represent, and further extends the model's generalization to novel capabilities such as generating non-English text spellings. We develop several small-scale expert plugins that efficiently endow a Stable Diffusion model with the capability to take different references. Each plugin is trained with auxiliary networks and loss functions customized for applications such as English scene-text generation, multi-lingual scene-text generation, and logo-image generation. Our expert plugins demonstrate superior results than the existing methods on all tasks, each containing only 28.55M trainable parameters.