Score
Designs, builds, and evaluates conditional generative models (e.g., conditional GANs) that synthesize three-dimensional outputs—such as voxel volumes, point clouds, or meshes—whose global or local attributes match specified target properties or labels. Work includes defining conditioning mechanisms, network architectures and loss functions, and evaluation metrics to ensure generated 3D structures satisfy requested quantitative properties while maintaining realism and diversity.
This paper presents a systematic survey of deep learning–driven 3D shape generation, addressing three core dimensions: shape representation, generative modeling, and evaluation protocols. Methodologically, it introduces the first unified taxonomy covering explicit (e.g., meshes), implicit (e.g., SDFs, NeRFs), and hybrid representations, traces the evolution of feedforward-based generative architectures, and consolidates major benchmarks (e.g., ShapeNet, FAUST) and metrics (e.g., Chamfer distance, Jensen–Shannon divergence), revealing inherent trade-offs among fidelity, diversity, and realism. Its principal contribution is a novel “representation–model–evaluation” triadic analytical framework, which explicitly identifies controllable shape modeling, efficient inference, and physically consistent generation as key open challenges. The framework establishes a structured benchmark and roadmap for future research in 3D generative modeling. (126 words)
Existing 3D generative models neglect scene-specific constraints, resulting in synthetic assets that fail to integrate naturally into real-world environments. To address this, we propose a scene-conditioned framework for 3D object style transfer and compositing. Our method jointly optimizes object texture and environment lighting via differentiable ray tracing, while leveraging image priors from pre-trained text-to-image diffusion models (e.g., Stable Diffusion) to ensure geometric–photometric consistency and semantic adaptability in an end-to-end manner. Crucially, this work establishes the first tight coupling between 3D stylization and 2D scene semantics, enabling dynamic, context-aware relighting and restyling of a single 3D object across diverse semantic settings (e.g., summer/winter, fantasy/futuristic). We validate our approach on varied indoor/outdoor scenes and arbitrary 3D objects, demonstrating substantial improvements in visual realism, physical plausibility, and artistic controllability of composited imagery.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
CAD modeling remains highly manual, lacking multimodal interaction and automation support. Method: This paper introduces the first end-to-end image-to-parametric-CAD-command-sequence framework for editable and manufacturable 3D shape generation. It innovatively integrates CLIP-style contrastive representation learning, latent diffusion priors, and an autoregressive Transformer architecture to enable image-driven CAD command sequence generation with geometric constraint-aware decoding. Contributions/Results: (1) Generates topologically valid, parameter-tunable, and manufacturing-ready CAD models from a single input image; (2) Enables cross-modal CAD retrieval, improving image-to-model accuracy by 32.7% on large-scale CAD databases; (3) Outperforms all state-of-the-art methods on both unconditional and image-conditioned CAD generation benchmarks. This work advances AI-driven design-to-manufacturing closed-loop automation.
Existing evaluation metrics for deep generative models (e.g., VAEs, GANs, diffusion models, Transformers) in engineering design—largely borrowed from statistical likelihood-based measures—fail to capture design-critical properties such as constraint satisfaction, functional performance, and design value. Method: We propose the first multidimensional evaluation framework tailored to engineering design, comprising four orthogonal dimensions: constraint compliance, functional effectiveness, novelty, and conditional controllability. We further develop an open-source, reproducible benchmark suite and software toolkit to bridge machine learning theory and design practice. Contribution/Results: The framework is rigorously validated on 2D visualization case studies and real-world engineering tasks—including bicycle frame and structural topology generation. Experiments demonstrate substantial improvements in alignment between automated evaluation and human-assessed design value: target achievement rate (+23.6%), geometric constraint compliance (+31.4%), and design novelty (+18.9%).
This work addresses three key challenges in structured image generation: (1) imprecise attribute control, (2) blurry outputs, and (3) oversimplified, unimodal prior modeling. To this end, we propose a multimodal disentangled generative framework based on conditional variational autoencoders (CVAEs). Methodologically, we introduce a weighted evidence lower bound (ELBO) optimization strategy that explicitly models multimodal priors over fine-grained attributes—such as hair color, eyewear presence, and species-specific traits—and enforces attribute-disentangled representations in the latent space. Notably, this is the first systematic application of CVAEs to attribute-driven, cross-domain structured generation (spanning faces and birds). Experiments on CelebA and CUB-200-2011 demonstrate substantial improvements in attribute accuracy, sample diversity, visual fidelity, and cross-category generalization, while maintaining robustness and interpretability.
This work addresses the challenge of evaluating conditional generation quality in compositional extrapolation settings, where the true target distribution is unavailable. The authors propose a post-hoc, instance-wise confidence scoring mechanism that requires no access to the target distribution. By constructing estimable metrics based on data manifold compatibility and attribute contrastive distance, the method holistically assesses both global realism and attribute fidelity. Notably, it incurs no additional training and is directly applicable to off-the-shelf pre-trained generative models. To the best of our knowledge, this is the first approach enabling effective evaluation of compositional extrapolation samples, facilitating sample filtering, ranking, and pre-generation abstention. Experiments on biological imaging and visual benchmarks demonstrate substantial improvements in morphological fidelity and downstream predictive performance, along with the capability for early abstention during generation.
This work addresses the challenge of reconstructing three-dimensional porous media with controllable porosity from only two-dimensional slice images, without reliance on expensive 3D training data. The authors propose a novel conditional generative adversarial network framework that uniquely integrates attribute-conditioned generation with 2D-to-3D reconstruction. Their approach employs a hybrid architecture featuring a 3D generator and a 2D discriminator, complemented by multi-axis slice extraction and an enhanced U-Net segmentation module. This method enables precise control over rock porosity while preserving 3D structural consistency, achieving strong performance on two types of carbonate rock samples: a porosity control coefficient of determination (R²) of 0.93, and mean absolute errors of 0.019 and 0.010 for heterogeneous and homogeneous samples, respectively.
Traditional inverse design methods are limited to point-wise target outputs and struggle to accommodate design requirements expressed as target distributions. This work formalizes, for the first time, the distribution-level inverse design problem and introduces a new paradigm termed Conditional Distribution Matching (CDM), defining two task variants: CDMS and CDMO. The authors propose MLGD-F, a plug-and-play inference algorithm that efficiently solves these tasks without additional training. MLGD-F leverages a pre-trained score-based diffusion model combined with a single-step conditional sampler, using a matching loss to guide gradient updates. The method successfully recovers inputs whose outputs align with complex target distributions—including discrete mixtures and continuous low-rank supports—demonstrating effectiveness across synthetic data, structured image transformation, and generative editing tasks.
This work addresses the absence of a unified classification framework for conditional 3D CT generation methods, which hinders systematic comparison and identification of critical design choices. To resolve this, the paper proposes a taxonomy centered on conditioning mechanisms, structured around three orthogonal dimensions: external knowledge type (Knowledge), knowledge integration paradigm (Integration), and generative architecture (Architecture), thereby defining a unified design space denoted as K×I×A. Through a comprehensive review of existing approaches, this framework not only clarifies prevailing technical pathways but also uncovers underexplored research directions. The resulting taxonomy offers theoretical guidance and a foundation for innovation in the design of future conditional 3D medical image generation methods.
Existing deep multivariate models are typically tailored to specific tasks, resulting in limited generalization capability. This work proposes a universal modeling framework that parameterizes the conditional distribution of each variable given all others using deep neural networks and represents the joint distribution through a Markov chain kernel. The model is trained by maximizing the likelihood under the stationary distribution of this kernel. By design, the approach eliminates the need for task-specific architectural modifications and inherently supports arbitrary downstream tasks as well as diverse semi-supervised learning scenarios. Consequently, it not only enhances model generalization but also significantly improves the efficiency of leveraging unlabeled data.