Score
Designs and trains probabilistic latent‑variable generative models that learn conditional distributions p(x|c) by encoding inputs into a latent space and decoding samples conditioned on c using a variational inference objective. These models are used to generate individual‑level conditional samples while capturing joint structure among attributes and preserving aggregate (population‑level) fidelity.
The rapid advancement of generative AI—including GANs, VAEs, and diffusion models—has led to an overwhelming and fragmented literature, necessitating a systematic synthesis. This survey proposes a unified technical taxonomy that integrates the evolutionary trajectories, architectural variants, and hybridization strategies of these three dominant paradigms, clarifying shared optimization principles for generation quality, diversity, and controllability. It introduces, for the first time, a multi-dimensional classification framework spanning model architecture, training mechanisms, and application domains. Furthermore, incorporating ethical considerations and societal impact, the survey identifies three key frontiers: scalability, trustworthy generation, and human-AI collaboration. By unifying conceptual foundations and highlighting emerging challenges, this work delivers a structured, forward-looking technical roadmap for researchers and practitioners in generative AI.
This work addresses the fragmentation between classical probabilistic latent variable models (PLVMs) and modern generative AI methods by unifying their underlying modeling principles, contrasting inference strategies, and analyzing representational trade-offs. We propose a unified probabilistic latent variable framework that systematically integrates seven canonical models: probabilistic PCA, hidden Markov models, variational autoencoders, normalizing flows, diffusion models, autoregressive models, and generative adversarial networks. Through formal analysis of their latent structures, inference mechanisms, and generative pathways, we construct the first theoretical taxonomy of generative AI. This framework clarifies the methodological evolution of generative modeling, strengthens its theoretical foundations, and provides interpretable conceptual guidance and structured design principles for developing novel architectures. (149 words)
This work addresses likelihood-free simulation-based inference (SBI), where the likelihood is intractable. To model complex, simulator-induced posterior distributions—often highly structured and multi-modal—we propose an efficient variational autoencoder (VAE)-based posterior estimation framework. Our method incorporates two complementary prior mechanisms: (1) a data-adaptive multivariate prior network to improve generalization across queries, and (2) a standard Gaussian prior to preserve model simplicity and expressiveness. End-to-end variational inference enables scalable, generative posterior approximation. On standard SBI benchmarks, our approach matches the accuracy of state-of-the-art normalizing flow–based methods while substantially reducing training and inference costs—yielding superior computational efficiency and scalability. The core contribution is the systematic integration of the VAE paradigm into SBI, establishing a lightweight, robust, and easily deployable solution for large-scale, high-dimensional simulation-based inference.
This work addresses the challenge of conditional sampling in generative diffusion models for Bayesian inverse problems. It systematically surveys and unifies two dominant paradigms: end-to-end methods based on the joint distribution, and decoupled approaches combining a pre-trained marginal distribution with an explicit likelihood model. We propose, for the first time, a theoretically consistent unified framework that integrates Monte Carlo sampling, diffusion process reweighting, conditional probability construction, and fine-tuning techniques—rigorously characterizing the underlying assumptions and intrinsic relationships among these methods. The framework bridges theoretical gaps across disparate conditional generation strategies and delivers a scalable, interpretable, and theoretically grounded toolkit for conditional sampling in scientific computing inverse problems, including image reconstruction and physics-based simulation.
Statistical analysis of black-box generative models—whose weights, pretraining data, and model covariates are inaccessible—remains challenging due to the absence of internal model information. Method: This paper introduces a data-centric kernel embedding framework that maps each generative model into a reproducing kernel Hilbert space (RKHS) induced by its output sample distribution, yielding model-level comparable representations. The method integrates functional-space projection, maximum mean discrepancy (MMD)-based distributional distance estimation, and nonparametric hypothesis testing to enable interpretable, cross-model statistical inference without requiring internal model access. Contribution/Results: It is the first approach to achieve purely input–output behavior-driven kernel-space embedding of generative models, circumventing black-box constraints. Evaluated on model clustering, anomaly detection, and performance attribution, it significantly outperforms baselines while exhibiting strong generalizability and plug-and-play applicability. This work establishes a novel, covariate-free paradigm for evaluating generative models under strict black-box conditions.
This work addresses conditional modeling of high-dimensional probability distributions by proposing a unified generative framework that jointly synthesizes conditional distributions across level sets of a collective variable ξ: ℝᵈ → ℝᵏ (k < d). To overcome inaccurate modeling caused by sparse sampling in low-probability level sets, we introduce an augmentation-driven data enrichment strategy enabling joint conditional density estimation across multiple level sets. The method integrates generative modeling, collective variable analysis, and conditional density estimation—constituting the first end-to-end approach for learning conditional distributions over multiple level sets simultaneously. Extensive numerical experiments demonstrate substantial improvements in generation fidelity and generalization performance within rare conformational regions. The framework establishes a novel paradigm for efficient data augmentation and high-precision conditional modeling, with direct applicability to molecular simulation and related computational science domains.
This work addresses a critical limitation in existing factorized generative models, which only match the marginal distribution of style latent variables without enforcing independence from class information, leading to conditional style leakage. The study demonstrates for the first time that marginal distribution matching alone is insufficient for effective disentanglement and establishes that four theoretical conditions must be jointly satisfied. To systematically quantify style-class leakage, the authors introduce a comprehensive auditing framework combining maximum mean discrepancy (MMD), linear probing, clustering evaluation, and multidimensional perturbation experiments. Empirical results reveal that multiple baseline models—despite achieving near-zero marginal MMD—still enable label recovery with 74%–100% accuracy. The proposed post-processing method substantially improves generation quality, attaining external evaluation scores of 0.97 on MNIST and 0.88 on CIFAR-10.
This work addresses the challenge of leveraging unlabeled data effectively in semi-supervised conditional generative modeling under label scarcity. The authors propose RepG, a framework that decouples the generation process into two stages: supervised sampling in a low-dimensional latent space and unsupervised reconstruction in the high-dimensional data space. By restricting conditional modeling to the low-dimensional space, RepG substantially reduces sample complexity and mitigates the curse of dimensionality. Theoretical analysis reveals that RepG achieves a faster non-asymptotic convergence rate through an error decomposition driven by conditional mutual information, and its optimality is confirmed by matching the minimax lower bound. Empirical results demonstrate the superior performance of RepG in semi-supervised conditional generation tasks.
This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.
This work addresses the lack of a general framework for modeling time-varying latent states in existing generative models, which often rely on auxiliary stochastic processes that are difficult to sample. The authors propose a novel approach that treats observation generation as a deterministic mapping of a tractable Markov process, employing an image-space stochastic process generator whose one-time marginal distribution matches that of a target projected process. The key innovation lies in extending Generator Matching—previously limited to static latent variables—to time-varying latent processes for the first time. By integrating stochastic process theory, Markov projections, and flow matching techniques, the method establishes a unified generative modeling framework. This framework not only subsumes existing models with discrete latent processes as special cases but also accommodates a broader class of time-varying latent conditions while rigorously ensuring consistency between the generated and target marginal distributions.
针对生成模型中先验与聚合后验不匹配的问题,提出使用聚合后验预测检查(APPC)方法,并通过实验验证了该方法的有效性。