Score
Designs, trains, and evaluates machine‑learned models that learn a data distribution to generate novel samples or conditional outputs. Work includes selecting generative architectures and training objectives (e.g., autoregressive, VAE, GAN, diffusion), implementing sampling and conditioning mechanisms, and measuring sample quality, diversity, and controllability.
The rapid advancement of generative AI—including GANs, VAEs, and diffusion models—has led to an overwhelming and fragmented literature, necessitating a systematic synthesis. This survey proposes a unified technical taxonomy that integrates the evolutionary trajectories, architectural variants, and hybridization strategies of these three dominant paradigms, clarifying shared optimization principles for generation quality, diversity, and controllability. It introduces, for the first time, a multi-dimensional classification framework spanning model architecture, training mechanisms, and application domains. Furthermore, incorporating ethical considerations and societal impact, the survey identifies three key frontiers: scalability, trustworthy generation, and human-AI collaboration. By unifying conceptual foundations and highlighting emerging challenges, this work delivers a structured, forward-looking technical roadmap for researchers and practitioners in generative AI.
This work addresses key bottlenecks in high-dimensional supervised learning and Bayesian inference—namely, strong parametric assumptions and reliance on large-scale, accurately labeled real-world data. We propose a model-agnostic generative modeling framework that leverages generative AI (Gen-AI) to synthesize high-fidelity training samples and employs deep neural networks for end-to-end nonparametric estimation of conditional densities and posterior quantiles—without assuming any prespecified distributional form. The framework unifies high-dimensional regression, dimensionality reduction (including feature selection), and uncertainty quantification. Experiments on the Ebola epidemic dataset demonstrate substantial improvements over conventional parametric methods in predictive accuracy, calibration, and computational scalability. To our knowledge, this is the first scalable, plug-and-play paradigm for model-free density estimation and Bayesian inference.
This work addresses the challenge of interpreting black-box machine learning models by proposing a generative diagnostic framework that reveals model data preferences and decision mechanisms through controllable synthetic samples. Methodologically, it formally defines and unifies three types of “model-query samples”—high-risk, parameter-sensitive, and model-comparative—via gradient-guided optimization, latent-space inversion, and constrained generative modeling, augmented by loss-sensitivity analysis for fine-grained semantic control. Extensive experiments across diverse architectures (CNNs, Transformers) and modalities (image, tabular data) demonstrate that the framework effectively characterizes decision boundaries, pinpoints vulnerability regions, and quantifies inter-model discrepancies. It significantly enhances the capability to verify model behavior interpretability, offering a general-purpose diagnostic tool applicable across tasks and model architectures.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
This paper addresses the conceptual and methodological challenges in bridging machine learning and computational creativity. It systematically traces the evolution of computational creativity theory and surveys key generative deep learning techniques—namely, Variational Autoencoders (VAEs), Generative Adversarial Networks (GANs), and Transformers—alongside their applications in creative tasks. To overcome persistent evaluation bottlenecks, the authors propose a hybrid assessment framework integrating cognitive modeling, multi-dimensional aesthetic metrics, and human-grounded benchmarks—marking the first comprehensive integration of theoretical paradigms with generative model practice. The core contributions are threefold: (1) construction of the most comprehensive research map of computational creativity to date; (2) a cross-paradigmatic methodological reflection on evaluation; and (3) identification of explainability enhancement and human-AI co-creation as critical frontiers—thereby establishing theoretical consensus and practical guidelines for algorithm design, evaluation standardization, and interdisciplinary deployment.
This paper challenges the reliance of machine learning research in the social sciences on abstract data-generating distributions, arguing that such assumptions lack empirical grounding in finite-population settings and engender interpretability and reproducibility issues. Method: The authors advocate replacing distributional assumptions with finite-population modeling, systematically advancing five core arguments grounded in statistical foundations, philosophical epistemology, and ML empirical analysis. They reconstruct the premises of learning theory by explicitly identifying the implicit assumptions and boundary conditions underlying distributional modeling. Contribution/Results: The proposed framework enhances theoretical coherence, modeling transparency, causal traceability, and practical applicability. It provides a novel paradigm and methodological foundation for sampling design, bias correction in evaluation, and reproducibility research—thereby addressing critical limitations of conventional distribution-based approaches in social-science ML applications.
Existing evaluation metrics for deep generative models (e.g., VAEs, GANs, diffusion models, Transformers) in engineering design—largely borrowed from statistical likelihood-based measures—fail to capture design-critical properties such as constraint satisfaction, functional performance, and design value. Method: We propose the first multidimensional evaluation framework tailored to engineering design, comprising four orthogonal dimensions: constraint compliance, functional effectiveness, novelty, and conditional controllability. We further develop an open-source, reproducible benchmark suite and software toolkit to bridge machine learning theory and design practice. Contribution/Results: The framework is rigorously validated on 2D visualization case studies and real-world engineering tasks—including bicycle frame and structural topology generation. Experiments demonstrate substantial improvements in alignment between automated evaluation and human-assessed design value: target achievement rate (+23.6%), geometric constraint compliance (+31.4%), and design novelty (+18.9%).
This study addresses the inverse design problem of gas turbine combustors by systematically evaluating the effectiveness of generative models in Bayesian inverse problems. For the first time in an engineering inverse design context, we benchmark conditional generative adversarial networks (cGAN), invertible neural networks (INN), conditional flow matching (CFM), and traditional Markov chain Monte Carlo (MCMC) methods, introducing a comprehensive evaluation metric that balances accuracy and diversity. Experimental results demonstrate that CFM significantly outperforms all other approaches across all metrics: it not only produces designs whose performance metrics align more closely with target specifications but also exhibits greater solution diversity and enhanced robustness to variations in training data size.
This study addresses the limitations of traditional Monte Carlo simulations, which rely on ad hoc assumptions and struggle to generate data reflecting realistic multilevel structures, thereby compromising the validity of quantitative method evaluations. To overcome this, the authors propose the first six-stage workflow integrating generative AI with multilevel data simulation. They innovatively adapt diffusion models and generative adversarial networks (GANs) to accommodate hierarchical data structures and introduce a comprehensive synthetic data quality assessment framework that ensures both within-table and cross-table consistency. Empirical experiments on real-world social science datasets demonstrate that the proposed approach substantially enhances the realism and reliability of Monte Carlo simulations, outperforming conventional strategies and providing a more empirically grounded benchmark for evaluating predictive performance and parameter recovery in quantitative methods.
This study investigates the evolutionary dynamics of generative models trained iteratively on synthetic data contaminated with real data, aiming to mitigate model collapse induced by data pollution. Through statistical modeling, mixture distribution analysis, and theoretical analysis of iterative training dynamics—complemented by theoretical derivations and simulations based on next-token prediction language models—the work demonstrates that model collapse can be effectively avoided and the true data distribution even recovered, provided the mixture weight of real data remains non-zero over time and is paired with sufficient sample sizes. This mechanism consistently enhances performance across diverse model classes, offering both theoretical guarantees and practical guidance for sustainable iterative training.
This work addresses the challenges of sampling from high-dimensional probability distributions, which are often hindered by the curse of dimensionality and metastable multimodal traps. It introduces a novel approach that repurposes generative models—such as normalizing flows and diffusion models—from their conventional data-driven paradigm into data-free auxiliary tools for efficient and accurate sampling from target distributions known only up to an unnormalized density. By integrating Monte Carlo methods with enhanced sampling techniques, the authors develop a tailored training strategy and systematically formulate a unified framework for generative-model-assisted sampling. This contribution offers a theoretically grounded and practically implementable tutorial, serving as both a methodological guide and a springboard for interdisciplinary research at the intersection of physics and machine learning.
This work addresses the lack of reliable statistical evaluation methods for generative models, which hinders the assessment of their generalization performance and the estimability of evaluation metrics from finite samples. The authors propose a theoretical framework that systematically analyzes the conditions under which common evaluation metrics are statistically estimable, distinguishing between test-class-based metrics and divergence-based metrics in finite-sample settings. Leveraging tools from integral probability metrics (IPMs), Rényi divergences, and fat-shattering dimension, they rigorously establish—for the first time—that IPMs induced by bounded test classes admit arbitrarily accurate estimation from finite samples, whereas KL and Rényi divergences, which depend on rare events, do not. This study provides a foundational theoretical basis and practical guidance for evaluating generative models.