text-to-image synthesis

Designs, implements, and evaluates generative systems that produce images from natural-language descriptions, encompassing text-guided or text-conditioned architectures (e.g., diffusion models, conditional generators), conditioning modules, training pipelines, and augmentation workflows. Measures and analyzes the fidelity of images to the input text, visual realism and diversity, and the usefulness of synthesized images for downstream objectives such as data augmentation or class-balancing.

text-to-imagesynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the current lack of systematic integration of image-generating generative AI in modeling and simulation. It presents the first comprehensive exploration of text-to-image generation techniques within this domain, proposing tool-agnostic, transferable principles and establishing a localized, reproducible generation pipeline that combines prompt engineering with simulation output mapping. The proposed approach supports diverse applications—including conceptual model representation, visualization of simulation results, generation of instructional materials, and construction of multi-scale model interfaces—thereby offering practitioners a structured knowledge framework to evaluate and adapt this emerging technology. By doing so, it significantly enhances the visual expressiveness and interactive capabilities of simulation systems.

generative AIimage generationmodeling and simulation

Controlled Training Data Generation with Diffusion Models

Mar 22, 2024
TY
Teresa Yeo
🏛️ Swiss Federal Institute of Technology Lausanne | MIT

This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.

Automate closed-loop feedback for adversarial prompt generationControl text-to-image models for supervised training dataGuide generation to match target data distributions

Nepotistically Trained Generative-AI Models Collapse

Nov 20, 2023
MB
Matyáš Boháček
🏛️ Stanford University | University of California, Berkeley

This work identifies a systematic degradation phenomenon—termed “nepotistic training”—occurring when generative AI models are fine-tuned using images synthesized by themselves. Leveraging diffusion-based text-to-image frameworks guided by CLIP, the study conducts controlled retraining experiments, multi-dimensional quality assessments (e.g., FID), and distributional shift analyses. It empirically demonstrates that the degradation is (i) contagious—injecting merely 0.5% AI-generated images degrades FID by over 200%; (ii) generalizable—distortions persist across unseen prompts; and (iii) irrecoverable—subsequent fine-tuning on clean real data fails to restore performance. The paper formally defines and validates these three core properties of this novel training failure mode. By establishing both theoretical insight and empirical evidence, the findings provide critical implications for sustainable generative model training and responsible content governance.

AI models distort images when retrained on their own outputsDistortion affects unrelated text prompts after retrainingModels fail to fully recover even with real data retraining

This work addresses the attribute-object-relation compositional misalignment problem in text-to-image generation, revealing structural deficiencies in the CLIP text encoder: its final-layer token embeddings are susceptible to interference from irrelevant words, and its attention mechanism struggles to capture fine-grained semantic relations. We propose a lightweight solution—fine-tuning only the CLIP linear projection head—without modifying the encoder backbone or the diffusion model architecture. On benchmarks such as Compositional COCO, our method improves compositional accuracy by 23.6% while preserving FID. Through attention reweighting and representational space analysis, we empirically demonstrate that suboptimal geometry in the CLIP embedding space is the primary cause of compositional failure. To our knowledge, this is the first systematic study establishing that projection-head fine-tuning alone suffices to substantially enhance compositional generalization—introducing a novel, efficient paradigm for improving multimodal alignment.

Addresses attribute binding failures in text-to-image modelsIdentifies sub-optimal CLIP text-encoder embeddings as a key issueProposes fine-tuning linear projection for improved compositional generation

Text-to-image diffusion models achieve high visual fidelity but struggle to faithfully render fine-grained subject details—such as precise text spelling—due to the limited semantic expressivity of standard text tokenizers. To address this, we propose a reference-guided generation framework based on lightweight expert plug-ins: leveraging a reference image as visual conditioning, jointly with textual input, to steer the Stable Diffusion architecture and overcome the representational bottlenecks of language-only conditioning. Our method integrates multiple task-specific, compact expert plug-ins (each with only 28.55M parameters), an auxiliary network, and dedicated loss functions tailored for cross-lingual text rendering and domain-specific applications like logo generation. Extensive experiments demonstrate consistent and significant improvements over existing state-of-the-art methods across English and multilingual text-to-image synthesis, as well as logo generation benchmarks.

Enhancing text-to-image models for precise subject renderingExtending diffusion models to generate multilingual text and logosOvercoming text tokenizer limitations with visual reference guidance

Latest Papers

What's happening recently
View more

Current generative models exhibit insufficient OCR capabilities in text-to-image generation and editing, particularly in producing legible, accurate, and layout-preserving textual content. Method: We propose the first systematic evaluation paradigm for “OCR-aware generation,” comprising 33 tasks across five real-world domains—document, handwritten, scene, artistic, and complex-layout images—and advocate photorealistic text generation as a foundational capability for general-purpose multimodal models. We introduce a customized input-prompt co-design mechanism and a multidimensional evaluation protocol integrating OCR accuracy metrics (e.g., CER, WER) with visual fidelity metrics (e.g., CLIP-Score, FID). Contribution/Results: Extensive benchmarking across six leading open- and closed-source models reveals critical deficiencies in character-level accuracy and structural layout preservation. Our analysis identifies language-vision alignment as the primary bottleneck limiting OCR generation performance, establishing a reproducible benchmark and concrete optimization directions for next-generation multimodal foundation models.

Assess OCR task performance across diverse text categoriesEvaluate generative models for text image generation and editingIdentify weaknesses in current models for photorealistic text generation

Although text-to-image (T2I) diffusion models from 2022 to 2025 have made remarkable progress in generating visually realistic and prompt-aligned images, their synthetic data consistently underperforms when used to train image classifiers. This study systematically evaluates the efficacy of data generated by successive generations of state-of-the-art T2I models through large-scale synthesis, standard classifier training protocols, and cross-model comparative analysis. The findings reveal that the pursuit of aesthetic quality has come at the cost of reduced data diversity and label consistency, leading to a disconnection between “generative realism” and “data utility.” Experiments demonstrate that classifiers trained on synthetic data from the latest T2I models exhibit significantly degraded accuracy on real-world test sets, indicating that current T2I-generated data is unsuitable as a reliable source for training robust image classifiers.

classification accuracydata realismsynthetic data

This work proposes a unified framework for understanding and developing generative artificial intelligence models capable of producing multimodal content, including images, text, video, and molecular structures. Addressing the current fragmentation in generative modeling, the study integrates core methodologies—such as variational autoencoders, generative adversarial networks, diffusion models, and large language models—into a cohesive theoretical system grounded in mathematical principles, architectural design, and mechanisms for controllable generation. This framework not only advances a systematic understanding of multimodal generative processes but also provides robust theoretical foundations and practical pathways for generating high-quality, controllable digital content, with direct implications for applications in scientific discovery and beyond.

ArchitecturesArtificial IntelligenceFoundational Principles

Policy Optimized Text-to-Image Pipeline Design

May 27, 2025
UG
Uri Gadot
🏛️ Technion | NVIDIA Research

Automated design of text-to-image multi-component pipelines faces two major bottlenecks: prohibitively high computational cost and poor generalization across tasks. Method: We propose the first end-to-end reinforcement learning framework that eliminates reliance on costly image generation for evaluation, introducing a novel image-generation-free ensemble reward model. Our approach employs a two-stage optimization strategy—lexical pretraining followed by Generalized Reinforcement Policy Optimization (GRPO)—and incorporates Classifier-Free Guidance (CFG)-guided model interpolation to enhance structural diversity. Contribution/Results: Without rendering any images, our method enables efficient workflow sequence modeling and search. It achieves state-of-the-art performance in image fidelity, structural novelty, and cross-task generalization, significantly outperforming existing baselines while reducing training overhead substantially.

Automating multi-component text-to-image pipeline design efficientlyImproving generalization and diversity in generated image workflowsReducing computational costs in pipeline optimization without image generation

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

Jun 06, 2025
TL
Tian Lan
🏛️ Beijing Institute of Technology | Peking University | Nanyang Technological University

Existing automated evaluation methods for generative content lack a systematic, cross-modal framework. Method: This paper conducts a large-scale literature review and cross-modal comparative analysis to establish, for the first time, a unified evaluation taxonomy covering text, image, and speech modalities. It identifies five fundamental evaluation paradigms and empirically validates their consistent applicability across three representative generative tasks. Furthermore, it introduces a comparability analysis framework to construct a structured knowledge graph that clarifies capability boundaries and limitations of existing methods per modality. Contributions/Results: (1) The first cross-modal unified classification system for generative evaluation; (2) abstraction of generalizable, transferable evaluation paradigms; and (3) a theoretical foundation and practical methodology for cross-modal consistent evaluation and joint metric design. This work bridges critical gaps in evaluating multimodal generative models and enables principled, interoperable assessment across modalities.

Identifying fundamental paradigms for cross-modal evaluation approachesLack of systematic framework for evaluating text, visual, and audio outputsNeed for unified taxonomy of automatic evaluation methods across modalities

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
KJ

K J Joseph

Research Scientist, Adobe Research
Deep Learning
JY

Jun-Yan Zhu

Assistant Professor, Carnegie Mellon University
Computer VisionComputer GraphicsGenerative ModelsComputational Photography
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
TN

Trung-Nghia Le

University of Science, VNU-HCM
Applied Deep LearningApplied Computer VisionMultimedia Security