Score
Designs, implements, and evaluates conditional generative models and hybrid generative pipelines that synthesize structured artifacts, simulations, or responses from explicit conditioning inputs (e.g., concept or semantic vectors, spectral targets, persona/profile signals, partial observations, or localized neighborhoods), using probabilistic formulations such as GANs, diffusion/score-based/ODE methods, and hybrid architectures. Builds training and sampling procedures, constraint- and fidelity-preserving conditioning mechanisms, localized ensemble or simulation workflows, and benchmarking metrics to measure sample quality, diversity, constraint satisfaction, and conditional accuracy.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
This work addresses the challenge of zero-shot conditional generation from pretrained unconditional diffusion models—specifically, generating samples satisfying complex logical constraints (e.g., structural conditions on tables, images, or time series) without fine-tuning. We propose a neural-symbolic soft-constraint embedding method that encodes first-order logic constraints as differentiable soft penalties and directly perturbs the score function to achieve theoretically consistent approximation of the conditional distribution—bypassing classifier-guided sampling or costly retraining. Our approach integrates score-based modeling, symbolic logic encoding, score correction, and stabilized sampling. Experiments across diverse data modalities demonstrate that our method achieves high-fidelity approximation of the true conditional distribution, significantly outperforming existing zero-shot conditional generation baselines.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
While generative models produce high-fidelity samples, they remain unreliable for compositional generalization—i.e., synthesizing novel concept combinations outside the training distribution. This work systematically investigates the compositional generalization of conditional diffusion models in a controlled synthetic setting. We decouple and independently manipulate three key data properties—frequency, structural composition, and factor disentanglement—to construct a multi-dimensional out-of-distribution (OOD) evaluation framework. Our core findings reveal a multiplicative emergence of compositional capability: overall performance equals the product of subtask accuracies; the order of capability emergence is governed by data generative structure; and low-frequency compositions incur substantially higher optimization costs. This is the first empirical demonstration of nonlinear, emergent compositional generalization in diffusion models, establishing quantitative links among data structure, capability emergence, and optimization cost—providing a data-centric theoretical foundation and design principles for trustworthy generative compositional reasoning.
This work addresses weak interpolation controllability and training instability in conditional generation. We propose Conditional Stochastic Interpolation (CSI), a framework that enables differentiable and controllable transport from a reference distribution to a target conditional distribution by modeling conditional probability flows or stochastic differential equations (SDEs). Key contributions include: (i) the first explicit formulation of the conditional drift and score function as conditional expectations; (ii) an adaptive diffusion term that enhances training stability; (iii) a non-asymptotic error bound guaranteeing convergence and generalization; and (iv) support for parameter-free regression estimation, deterministic ODE sampling, and adaptive-diffusion sampling. Extensive experiments on standard image datasets demonstrate high-quality, high-fidelity conditional generation. The method combines theoretical rigor—grounded in conditional probability flow theory—with practical effectiveness, offering improved controllability, stability, and flexibility over existing approaches.
This study addresses the inverse design problem of gas turbine combustors by systematically evaluating the effectiveness of generative models in Bayesian inverse problems. For the first time in an engineering inverse design context, we benchmark conditional generative adversarial networks (cGAN), invertible neural networks (INN), conditional flow matching (CFM), and traditional Markov chain Monte Carlo (MCMC) methods, introducing a comprehensive evaluation metric that balances accuracy and diversity. Experimental results demonstrate that CFM significantly outperforms all other approaches across all metrics: it not only produces designs whose performance metrics align more closely with target specifications but also exhibits greater solution diversity and enhanced robustness to variations in training data size.
This work addresses the challenge of evaluating conditional generation quality in compositional extrapolation settings, where the true target distribution is unavailable. The authors propose a post-hoc, instance-wise confidence scoring mechanism that requires no access to the target distribution. By constructing estimable metrics based on data manifold compatibility and attribute contrastive distance, the method holistically assesses both global realism and attribute fidelity. Notably, it incurs no additional training and is directly applicable to off-the-shelf pre-trained generative models. To the best of our knowledge, this is the first approach enabling effective evaluation of compositional extrapolation samples, facilitating sample filtering, ranking, and pre-generation abstention. Experiments on biological imaging and visual benchmarks demonstrate substantial improvements in morphological fidelity and downstream predictive performance, along with the capability for early abstention during generation.
This work addresses the limitation of existing generative models, which rely on fixed pre-trained parameters and lack the ability to dynamically adapt to individual input instances. To overcome this, the authors propose Composer, a framework that enables instance-level adaptation at test time by conditionally generating and injecting lightweight parameters based on the input, without requiring fine-tuning. Composer introduces, for the first time, a test-time mechanism for instance-specific parameter composition, endowing static models with context-awareness and dynamic adaptability. The approach is compatible with both diffusion and autoregressive architectures and supports quantized deployment. Experimental results demonstrate that Composer consistently enhances generation quality across diverse tasks while maintaining low computational and memory overhead, confirming its effectiveness and broad applicability.
This work addresses the challenge of precise control in conditional discrete generative models when confronted with unseen condition combinations. The authors propose a theory-driven, composable discrete generation framework that integrates parallel token prediction with an absorbing diffusion mechanism and a concept-weighted conditional fusion strategy. This approach enables accurate modeling of an arbitrary number and combination of conditions while unifying mask-based generation within the same paradigm. Leveraging compositional vocabularies derived from VQ-VAE/VQ-GAN, the method achieves a 63.4% average reduction in error rate, a 9.58 improvement in FID, and 2.3–12× faster inference across three datasets. Furthermore, it successfully extends to pretrained text-to-image models, enabling fine-grained controllable generation.
This work addresses the performance instability of conditional generative models across diverse tasks or inputs by proposing a novel model averaging framework that requires no explicit density estimation and relies solely on conditional samples. The approach leverages Maximum Mean Discrepancy (MMD) to measure distances between conditional distributions and introduces two fusion mechanisms: static weighting (StaticMA) and input-adaptive weighting (MoEMA). Notably, this study provides the first theoretical guarantees for MoEMA, establishing its asymptotic optimality and consistency of the learned weighting function. By incorporating softmax-gated neural networks and representation mappings, the method generalizes effectively to unstructured data. Empirical evaluations demonstrate that MoEMA consistently outperforms existing baselines across multimodal domains, including tabular, image, and text data.