Decoder-Only LLMs are Better Controllers for Diffusion Models

📅 2025-02-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current text-to-image diffusion models suffer from limited semantic representation capacity of pretrained text encoders, resulting in poor prompt robustness, weak text-image alignment, and labor-intensive hyperparameter tuning. To address these issues, we propose a lightweight adapter framework that systematically demonstrates and validates decoder-only large language models (LLMs) as superior semantic controllers for diffusion models compared to encoder-only or encoder-decoder architectures. Our method employs learnable adapters to efficiently bridge LLMs with diffusion backbones and introduces cross-modal feature alignment alongside theoretical analysis—specifically examining attention mechanisms and representation capacity. Extensive experiments across multiple benchmarks show substantial improvements over state-of-the-art methods, achieving significant gains in text-image alignment fidelity, fine-grained detail accuracy, and prompt robustness, while markedly reducing reliance on manual intervention.

Technology Category

Computer Vision: Diffusion Models for VisionMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
Groundbreaking advancements in text-to-image generation have recently been achieved with the emergence of diffusion models. These models exhibit a remarkable ability to generate highly artistic and intricately detailed images based on textual prompts. However, obtaining desired generation outcomes often necessitates repetitive trials of manipulating text prompts just like casting spells on a magic mirror, and the reason behind that is the limited capability of semantic understanding inherent in current image generation models. Specifically, existing diffusion models encode the text prompt input with a pre-trained encoder structure, which is usually trained on a limited number of image-caption pairs. The state-of-the-art large language models (LLMs) based on the decoder-only structure have shown a powerful semantic understanding capability as their architectures are more suitable for training on very large-scale unlabeled data. In this work, we propose to enhance text-to-image diffusion models by borrowing the strength of semantic understanding from large language models, and devise a simple yet effective adapter to allow the diffusion models to be compatible with the decoder-only structure. Meanwhile, we also provide a supporting theoretical analysis with various architectures (e.g., encoder-only, encoder-decoder, and decoder-only), and conduct extensive empirical evaluations to verify its effectiveness. The experimental results show that the enhanced models with our adapter module are superior to the stat-of-the-art models in terms of text-to-image generation quality and reliability.
Problem

Research questions and friction points this paper is trying to address.

Enhance text-to-image diffusion models' semantic understanding.
Compatibility of diffusion models with decoder-only LLMs.
Improve text-to-image generation quality and reliability.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoder-Only LLMs enhance diffusion models
Adapter bridges LLMs and diffusion models
Improved text-to-image generation quality