Score
Design and implement generative models that take a scene or image as conditioning input and produce diverse conditional outputs—such as images, latent goal samples, or predicted future states—by constructing conditional latent spaces and sampling mechanisms. Analyze training objectives, model architectures, and evaluation procedures to capture multimodal goal distributions, incorporate contextual signals (e.g., human pose), and generalize to novel scenes without relying on scene labels.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work systematically evaluates the applicability of generative AI to scientific image understanding, focusing on text-to-image and image-to-image generation tasks. We propose the first horizontal evaluation framework tailored to scientific imaging scenarios, benchmarking three dominant generative architectures—VAEs, GANs, and diffusion models—across six quantitative dimensions: fidelity, controllability, physical consistency, noise robustness, fine-grained detail accuracy, and domain adaptation efficiency. To address domain-specific requirements, we introduce novel evaluation metrics for generative quality in scientific imaging. Our analysis reveals fundamental trade-offs among key performance indicators across architectures. Furthermore, we identify concrete technical pathways toward enhancing model interpretability. Collectively, these findings provide both theoretical foundations and practical guidelines for the reliable deployment of generative AI in computational imaging, microscopy analysis, and other scientific domains.
Existing generative models lack robust scene understanding—particularly in handling variable object counts, diverse shapes, and distribution shifts relative to training data. Method: This work frames visual understanding as inverse inference over compositional generative models. We propose the first compositional inverse generation framework: it performs unsupervised inversion via energy-based modeling, modularly assembles generative components, and enables zero-shot adaptation of pretrained text-to-image models (e.g., Stable Diffusion) without fine-tuning. Contribution/Results: The method achieves multi-object perception without parameter updates, demonstrating strong generalization to unseen object counts, geometric configurations, and scene distributions. It accurately decomposes scenes into constituent objects and global scene factors—even in novel environments—thereby significantly improving robustness in multi-object recognition and enhancing structural interpretability of generated explanations.
This paper addresses key deployment bottlenecks hindering practical adoption of generative AI for image synthesis—namely, high computational overhead, data bias, and poor alignment with user intent. To tackle these challenges, we propose a structured, input-modality–centric taxonomy that unifies modeling across GANs, diffusion models, and conditional generation paradigms. We systematically categorize core tasks—including image-to-image translation, text-to-image generation, domain adaptation, and multimodal alignment—and conduct an in-depth analysis of architectural design principles and applicability boundaries of representative models such as DALL·E, ControlNet, and DeepSeek Janus-Pro. Furthermore, we establish an industrial-deployment–oriented evaluation framework, explicitly delineating optimization pathways for computational efficiency, bias mitigation, and intent alignment. The resulting methodology provides researchers and practitioners with a theoretically grounded yet practically actionable guide for developing and deploying robust, equitable, and controllable generative image systems.
Existing 3D generative models neglect scene-specific constraints, resulting in synthetic assets that fail to integrate naturally into real-world environments. To address this, we propose a scene-conditioned framework for 3D object style transfer and compositing. Our method jointly optimizes object texture and environment lighting via differentiable ray tracing, while leveraging image priors from pre-trained text-to-image diffusion models (e.g., Stable Diffusion) to ensure geometric–photometric consistency and semantic adaptability in an end-to-end manner. Crucially, this work establishes the first tight coupling between 3D stylization and 2D scene semantics, enabling dynamic, context-aware relighting and restyling of a single 3D object across diverse semantic settings (e.g., summer/winter, fantasy/futuristic). We validate our approach on varied indoor/outdoor scenes and arbitrary 3D objects, demonstrating substantial improvements in visual realism, physical plausibility, and artistic controllability of composited imagery.
This work addresses three key challenges in structured image generation: (1) imprecise attribute control, (2) blurry outputs, and (3) oversimplified, unimodal prior modeling. To this end, we propose a multimodal disentangled generative framework based on conditional variational autoencoders (CVAEs). Methodologically, we introduce a weighted evidence lower bound (ELBO) optimization strategy that explicitly models multimodal priors over fine-grained attributes—such as hair color, eyewear presence, and species-specific traits—and enforces attribute-disentangled representations in the latent space. Notably, this is the first systematic application of CVAEs to attribute-driven, cross-domain structured generation (spanning faces and birds). Experiments on CelebA and CUB-200-2011 demonstrate substantial improvements in attribute accuracy, sample diversity, visual fidelity, and cross-category generalization, while maintaining robustness and interpretability.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
This work proposes a unified framework for understanding and developing generative artificial intelligence models capable of producing multimodal content, including images, text, video, and molecular structures. Addressing the current fragmentation in generative modeling, the study integrates core methodologies—such as variational autoencoders, generative adversarial networks, diffusion models, and large language models—into a cohesive theoretical system grounded in mathematical principles, architectural design, and mechanisms for controllable generation. This framework not only advances a systematic understanding of multimodal generative processes but also provides robust theoretical foundations and practical pathways for generating high-quality, controllable digital content, with direct implications for applications in scientific discovery and beyond.
This work investigates whether generative modeling is structurally necessary for data-efficient, human-like visual perception—particularly compositional generalization. Theoretically, we establish for the first time that, under compositional data-generating mechanisms, generative approaches inherently encode critical inductive biases via decoder constraints and inverse inference, whereas discriminative methods cannot replicate these biases equivalently through regularization or architectural design alone. Methodologically, we integrate gradient-based online search, generative replay, decoder-constrained modeling, and theory-driven inductive bias analysis. Experiments on photorealistic image datasets demonstrate that—even with minimal decoder design—generative models achieve substantial gains in compositional generalization, without requiring additional data, large-scale pretraining, or supervised signals. These results empirically validate the structural necessity of generative capacity for human-like visual generalization.
Existing deep multivariate models are typically tailored to specific tasks, resulting in limited generalization capability. This work proposes a universal modeling framework that parameterizes the conditional distribution of each variable given all others using deep neural networks and represents the joint distribution through a Markov chain kernel. The model is trained by maximizing the likelihood under the stationary distribution of this kernel. By design, the approach eliminates the need for task-specific architectural modifications and inherently supports arbitrary downstream tasks as well as diverse semi-supervised learning scenarios. Consequently, it not only enhances model generalization but also significantly improves the efficiency of leveraging unlabeled data.
This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.
This work addresses the lack of a general framework for modeling time-varying latent states in existing generative models, which often rely on auxiliary stochastic processes that are difficult to sample. The authors propose a novel approach that treats observation generation as a deterministic mapping of a tractable Markov process, employing an image-space stochastic process generator whose one-time marginal distribution matches that of a target projected process. The key innovation lies in extending Generator Matching—previously limited to static latent variables—to time-varying latent processes for the first time. By integrating stochastic process theory, Markov projections, and flow matching techniques, the method establishes a unified generative modeling framework. This framework not only subsumes existing models with discrete latent processes as special cases but also accommodates a broader class of time-varying latent conditions while rigorously ensuring consistency between the generated and target marginal distributions.