image generation

Designs and builds algorithms, models, and pipelines that synthesize visual images (图像生成) from inputs such as random noise, text prompts, condition maps, or other media. This competence covers training and evaluating generative models, designing loss functions and architectures, curating and preprocessing image datasets, and implementing inference and rendering workflows plus metrics for image quality, diversity, and fidelity.

imagegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.53
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the current lack of systematic integration of image-generating generative AI in modeling and simulation. It presents the first comprehensive exploration of text-to-image generation techniques within this domain, proposing tool-agnostic, transferable principles and establishing a localized, reproducible generation pipeline that combines prompt engineering with simulation output mapping. The proposed approach supports diverse applications—including conceptual model representation, visualization of simulation results, generation of instructional materials, and construction of multi-scale model interfaces—thereby offering practitioners a structured knowledge framework to evaluate and adapt this emerging technology. By doing so, it significantly enhances the visual expressiveness and interactive capabilities of simulation systems.

generative AIimage generationmodeling and simulation

An Empirical Study of GPT-4o Image Generation Capabilities

Apr 08, 2025
SC
Sixiang Chen
🏛️ The Hong Kong University of Science and Technology | National University of Singapore | Peking University | The Chinese University of Hong Kong | University of Washington | Wuhan University

Prior work lacks systematic evaluation of unified multimodal generation capabilities in state-of-the-art large multimodal models (LMMs) across diverse cross-modal image generation tasks (e.g., text-to-image, image-to-image, image-to-3D, image-to-X). Method: This study conducts the first comprehensive benchmark of GPT-4o across 20+ cross-modal generation tasks, employing multi-task prompt engineering, cross-modal consistency assessment, and a hybrid human–automated evaluation framework integrating quantitative metrics and qualitative analysis. Contribution/Results: GPT-4o demonstrates superior text–image co-generation performance versus mainstream multimodal models, reflecting strong semantic understanding and cross-modal alignment. However, it lags significantly behind specialized diffusion models in fine-grained spatial control, 3D geometric fidelity, and complex image editing—positioning its overall generative capability between domain-specific models and earlier LMMs. The study empirically identifies data scale and architectural design as critical determinants of generation quality under unified architectures, establishing the first evidence-based benchmark and methodological paradigm for evaluating generative capabilities of multimodal foundation models.

Assessing GPT-4o's image generation capabilities empiricallyComparing GPT-4o with open-source and commercial modelsExploring unified frameworks for text and image generation

Generative AI for Vision: A Comprehensive Study of Frameworks and Applications

Jan 29, 2025
FB
Fouad Bousetouane
🏛️ The University of Chicago | 2ndsight.ai

This paper addresses key deployment bottlenecks hindering practical adoption of generative AI for image synthesis—namely, high computational overhead, data bias, and poor alignment with user intent. To tackle these challenges, we propose a structured, input-modality–centric taxonomy that unifies modeling across GANs, diffusion models, and conditional generation paradigms. We systematically categorize core tasks—including image-to-image translation, text-to-image generation, domain adaptation, and multimodal alignment—and conduct an in-depth analysis of architectural design principles and applicability boundaries of representative models such as DALL·E, ControlNet, and DeepSeek Janus-Pro. Furthermore, we establish an industrial-deployment–oriented evaluation framework, explicitly delineating optimization pathways for computational efficiency, bias mitigation, and intent alignment. The resulting methodology provides researchers and practitioners with a theoretically grounded yet practically actionable guide for developing and deploying robust, equitable, and controllable generative image systems.

Artificial IntelligenceComputational CostImage Generation

Nepotistically Trained Generative-AI Models Collapse

Nov 20, 2023
MB
Matyáš Boháček
🏛️ Stanford University | University of California, Berkeley

This work identifies a systematic degradation phenomenon—termed “nepotistic training”—occurring when generative AI models are fine-tuned using images synthesized by themselves. Leveraging diffusion-based text-to-image frameworks guided by CLIP, the study conducts controlled retraining experiments, multi-dimensional quality assessments (e.g., FID), and distributional shift analyses. It empirically demonstrates that the degradation is (i) contagious—injecting merely 0.5% AI-generated images degrades FID by over 200%; (ii) generalizable—distortions persist across unseen prompts; and (iii) irrecoverable—subsequent fine-tuning on clean real data fails to restore performance. The paper formally defines and validates these three core properties of this novel training failure mode. By establishing both theoretical insight and empirical evidence, the findings provide critical implications for sustainable generative model training and responsible content governance.

AI models distort images when retrained on their own outputsDistortion affects unrelated text prompts after retrainingModels fail to fully recover even with real data retraining

Latest Papers

What's happening recently
View more

This work presents the first comprehensive survey of mainstream image generation techniques developed over the past decade, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, autoregressive models, Transformers, and diffusion models. It systematically traces their evolution in terms of objective functions, architectural designs, and training algorithms, while also extending the discussion to applications such as video generation, deepfake detection, and watermarking. By constructing a clear technological roadmap, the paper synthesizes optimization strategies, failure modes, and inherent limitations across model families, highlighting the progression from static image synthesis to high-quality video generation. Furthermore, it underscores critical ethical and safety considerations regarding model robustness and responsible deployment, offering a systematic reference for future research and practical applications in generative modeling.

generative modelsimage generationliterature fragmentation

Existing unified multimodal models struggle to emulate the human-like capacity for iterative reasoning and progressive refinement over intermediate visual states during drawing. This work proposes a process-driven image generation paradigm that models synthesis as a multi-round, interleaved reasoning loop of “text planning → visual sketching → text reflection → visual refinement.” It introduces, for the first time, a multi-step inference mechanism with bidirectional constraints between textual instructions and visual states, enabling dynamic evaluation and correction of intermediate outputs. Through dense step-wise supervision, spatial-semantic consistency constraints, and strategies to preserve textual priors, the approach ensures interpretability, controllability, and plausibility of intermediate states throughout generation. Experiments demonstrate significant improvements in structural coherence and detail fidelity across multiple text-to-image benchmarks.

intermediate statesmultimodal reasoningprocess-driven image generation

This work addresses the disconnect between sampling and prototyping in traditional creative workflows, which hinders rapid exploration and problem framing in the generative AI era. To bridge this gap, the paper introduces “protosampling”—a novel paradigm that unifies sampling and prototyping within a cohesive creative practice. Implemented through Atelier, an interactive canvas system, this approach integrates multimodal generative models, intelligent search, and visual collection management to enable users to instantly generate, blend, and organize visual content within a shared workspace. Atelier facilitates dynamic ideation by allowing creative practitioners to synthesize fragmented concepts fluidly, thereby supporting a seamless transition from divergent thinking to the consolidation of actionable design solutions.

creative processgenerative AIproblem construction

This study investigates whether current text-to-image generation models can preserve mathematical correctness when producing visual representations—such as diagrams or geometric constructions—required to solve mathematical problems. To this end, the authors construct a benchmark of 900 tasks spanning seven core mathematical domains and introduce, for the first time, an automated evaluation framework tailored to mathematical visual generation. This framework combines executable verifiers with a Script-as-a-Judge protocol to enable objective assessment. Experimental results reveal that even the best closed-source model achieves only 42.0% overall accuracy, while open-source models generally score below 11%, approaching 0% on structured tasks. These findings demonstrate that existing text-to-image models fundamentally lack the capability to generate mathematically valid visual content.

generative modelsmathematical competencemathematical fidelity

This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.

generative explorationimage generationnon-literal visual combinations

Hot Scholars

XB

Xiang Bai

Huazhong University of Science and Technology (HUST)
Computer VisionOCR
VM

Vishal M. Patel

Associate Professor, ECE, Johns Hopkins University
Image ProcessingComputer VisionBiometricsMedical Image Analysis
MD

Mingyu Ding

Assistant Professor, UNC Chapel Hill
RoboticsEmbodied AIComputer Vision
TD

Trevor Darrell

Professor of Computer Science, U.C. Berkeley
Computer VisionArtificial IntelligenceAIMachine Learning
YW

Yongliang Wu

Southeast University
Vision-Language Model