Score
Designs, builds, and evaluates mechanisms that adaptively modulate conditioning inputs to conditional generative models—e.g., dynamic weighting of auxiliary networks, prompt embeddings, or visual prompts—to enforce structural constraints while preserving the model’s expressive diversity and realism. Includes algorithms for scheduling or blending conditioning signals, detecting and suppressing over‑conditioning artifacts, and formulating conditioning strategies that balance constraint adherence with generative quality.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work addresses the limitation of existing generative models, which rely on fixed pre-trained parameters and lack the ability to dynamically adapt to individual input instances. To overcome this, the authors propose Composer, a framework that enables instance-level adaptation at test time by conditionally generating and injecting lightweight parameters based on the input, without requiring fine-tuning. Composer introduces, for the first time, a test-time mechanism for instance-specific parameter composition, endowing static models with context-awareness and dynamic adaptability. The approach is compatible with both diffusion and autoregressive architectures and supports quantized deployment. Experimental results demonstrate that Composer consistently enhances generation quality across diverse tasks while maintaining low computational and memory overhead, confirming its effectiveness and broad applicability.
This work addresses the challenge of zero-shot conditional generation from pretrained unconditional diffusion models—specifically, generating samples satisfying complex logical constraints (e.g., structural conditions on tables, images, or time series) without fine-tuning. We propose a neural-symbolic soft-constraint embedding method that encodes first-order logic constraints as differentiable soft penalties and directly perturbs the score function to achieve theoretically consistent approximation of the conditional distribution—bypassing classifier-guided sampling or costly retraining. Our approach integrates score-based modeling, symbolic logic encoding, score correction, and stabilized sampling. Experiments across diverse data modalities demonstrate that our method achieves high-fidelity approximation of the true conditional distribution, significantly outperforming existing zero-shot conditional generation baselines.
To address the challenges of few-shot learning, sparse labeling, heterogeneous (numerical and categorical) conditioning variables, and constrained computational resources in engineering applications, this paper proposes a masked conditional generative paradigm. We design a unified learnable embedding to jointly model heterogeneous conditions and introduce a masked conditional scheduling mechanism that explicitly simulates missing conditions during training to enhance robustness to incomplete inputs. Furthermore, we construct a lightweight collaborative architecture integrating a variational autoencoder and a latent diffusion model, coupled with knowledge distillation from pre-trained large models. Experiments on 2D point cloud and engineering image datasets demonstrate that the method enables efficient training with only a small number of labeled samples; achieves a 32% reduction in Fréchet Inception Distance (FID); significantly improves conditional fidelity; and simultaneously ensures strong controllability and high generation quality.
Under rapid generative AI model iteration, users’ ability to adapt to evolving models critically determines the translation of technological advancement into economic value. Method: This paper introduces *prompt adaptation*—users’ deliberate refinement of input prompts—as a dynamic complementarity mechanism in generative AI. We systematically quantify its impact via online controlled experiments, analysis of over 18,000 real-world human-AI interactions, and image similarity-based evaluation. Contribution/Results: Approximately 49% of DALL·E 3’s performance gain stems from user-initiated prompt adjustments; replacing such human adaptation with automated prompt rewriting incurs a 58% loss in upgrade benefits. These findings demonstrate that user adaptive behavior accounts for nearly half of the value generated by model upgrades, underscoring its pivotal role in human-AI co-creation. The study provides empirical grounding for AI deployment strategies and human-centered interface design.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
The compositional generalization capability of visual generative models—i.e., their ability to synthesize novel combinations of known concepts—remains poorly understood. This paper systematically investigates key architectural and training factors affecting compositional generalization in image and video generation, identifying two core drivers: (1) the discrete versus continuous nature of the training objective, and (2) the information completeness of conditional inputs regarding concept composition. To address these, we propose a hybrid optimization strategy within the MaskGIT framework: augmenting the primary discrete reconstruction objective with an auxiliary continuous target derived from Joint-Embedding Predictive Architecture (JEPA), thereby relaxing the discrete loss. We conduct controlled ablation studies to quantitatively evaluate compositional generalization. Experiments demonstrate substantial improvements in compositional generalization for discrete generative models on complex scenes. To our knowledge, this is the first work to empirically validate the effectiveness and generality of jointly optimizing discrete and continuous objectives for structured semantic synthesis.
This work addresses the limited interpretability and controllability of image generation models by proposing an internal mechanism intervention method based on parameterized activation functions. Specifically, we replace standard activations (e.g., ReLU) in mainstream generative architectures—such as StyleGAN2 and BigGAN—with learnable, semantically interpretable parameterized variants (e.g., generalized Swish with shape and bias controls). This enables direct, fine-grained manipulation of activation behavior for targeted image editing, without altering network architecture or requiring additional training. We demonstrate effective, attribute-specific control—including illumination, texture, and pose—on FFHQ and ImageNet. Experimental results confirm that our intervention preserves model fidelity while offering both human-understandable semantics and quantitative effectiveness. The approach establishes a novel paradigm for transparent, plug-and-play control over generative models’ internal representations.
This work addresses the challenge of enforcing complex nonlinear constraints—such as road-legal regions in robotic control and autonomous driving—within generative models, where existing approaches often fail to simultaneously ensure constraint satisfaction and high-fidelity generation. The authors propose a constrained fine-tuning framework that leverages pre-trained generative models to produce outputs strictly confined within structured feasible regions, without compromising sample realism. By overcoming the limitations of conventional fine-tuning or training-free strategies, the method achieves superior performance across diverse and intricate constraint scenarios, consistently outperforming current baselines in both generation quality and adherence to constraints.
Existing image synthesis methods often struggle to simultaneously achieve 3D structural consistency and high photorealism, frequently introducing over-constrained artifacts due to excessive reliance on visual cues. This work proposes an adaptive conditioning framework based on diffusion models that dynamically modulates ControlNet guidance strength through a self-supervised mechanism, thereby mitigating over-constraint while preserving generative expressiveness. Additionally, a multi-agent vision-language model is introduced to produce semantically rich textual prompts aligned with 3D geometry. The proposed approach significantly enhances both visual fidelity and 3D consistency of synthesized images, enabling the generation of high-quality, scalable datasets with precise 2D/3D annotations, which demonstrate superior utility in downstream tasks.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.