generative model conditioning

Designs, implements, and evaluates conditioning mechanisms for generative models by integrating auxiliary conditioning networks (e.g., ControlNet) or adapters that accept structured control signals and route them into a model's training and inference pipelines. This includes building preprocessing and encoding of control inputs, selecting injection points (cross‑attention, feature fusion, etc.), training or fine‑tuning the conditioning modules, and measuring how conditioning alters or constrains generated outputs.

generativemodelconditioning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Masked Conditioning for Deep Generative Models

May 22, 2025
PM
Phillip Mueller
🏛️ BMW Group | University of Augsburg | Ludwig-Maximilians-University Munich

To address the challenges of few-shot learning, sparse labeling, heterogeneous (numerical and categorical) conditioning variables, and constrained computational resources in engineering applications, this paper proposes a masked conditional generative paradigm. We design a unified learnable embedding to jointly model heterogeneous conditions and introduce a masked conditional scheduling mechanism that explicitly simulates missing conditions during training to enhance robustness to incomplete inputs. Furthermore, we construct a lightweight collaborative architecture integrating a variational autoencoder and a latent diffusion model, coupled with knowledge distillation from pre-trained large models. Experiments on 2D point cloud and engineering image datasets demonstrate that the method enables efficient training with only a small number of labeled samples; achieves a 32% reduction in Fréchet Inception Distance (FID); significantly improves conditional fidelity; and simultaneously ensures strong controllability and high generation quality.

Enabling generative models with limited computational resourcesHandling small, sparse, mixed-type datasets in engineeringImproving generation quality with small models and pretrained foundations

Controlling the image generation process with parametric activation functions

Oct 17, 2025
IP
Ilia Pavlov
🏛️ University of the Arts London

This work addresses the limited interpretability and controllability of image generation models by proposing an internal mechanism intervention method based on parameterized activation functions. Specifically, we replace standard activations (e.g., ReLU) in mainstream generative architectures—such as StyleGAN2 and BigGAN—with learnable, semantically interpretable parameterized variants (e.g., generalized Swish with shape and bias controls). This enables direct, fine-grained manipulation of activation behavior for targeted image editing, without altering network architecture or requiring additional training. We demonstrate effective, attribute-specific control—including illumination, texture, and pose—on FFHQ and ImageNet. Experimental results confirm that our intervention preserves model fidelity while offering both human-understandable semantics and quantitative effectiveness. The approach establishes a novel paradigm for transparent, plug-and-play control over generative models’ internal representations.

Demonstrating method effectiveness on StyleGAN2 and BigGAN networksDeveloping interactive tools for interpretable control of generative modelsReplacing activation functions with parametric alternatives for output manipulation

Automatically Adaptive Conformal Risk Control

Jun 25, 2024
VB
Vincent Blot
🏛️ Paris-Saclay University | Capgemini Invent France | University of California, Berkeley | INRIA Paris | ENSIIE

For black-box models applied to complex tasks such as image segmentation, defining meaningful conditional events is challenging, leading to uncertainty estimates that fail to reflect inherent sample difficulty. Method: This paper proposes an input-dependent statistical risk control framework grounded in conformal prediction. It introduces a novel, algorithm-driven mechanism for dynamically selecting conditional function classes—bypassing manual discretization—by adaptively constructing these classes based on test-sample difficulty and integrating online parameter tuning for fine-grained, approximately conditional risk control. Contribution/Results: Experiments on regression and image segmentation demonstrate substantial improvements in uncertainty calibration accuracy. The method guarantees strict statistical risk control while enhancing generalization robustness and predictive reliability.

Achieves conditional control of statistical risks adaptively.Enables finer-grained control over model performance in regression and segmentation.Ensures reliable performance of black-box machine learning algorithms.

Amortized Probabilistic Conditioning for Optimization, Simulation and Inference

Oct 20, 2024
PE
Paul E. Chang
🏛️ University of Helsinki | Aalto University | University of Manchester

Existing amortized meta-learning approaches struggle to flexibly condition on and extract probabilistic latent variables at inference time, limiting their applicability in Bayesian inference, optimization, and simulation. To address this, we propose the Amortized Conditioning Engine (ACE), a Transformer-based, conditional neural process model that enables *bidirectional explicit manipulation* of probabilistic latent variables for the first time—supporting dynamic injection of observed data and interpretable latents, prior embedding, and joint generation of discrete/continuous data and predictive latent distributions. ACE unifies variational inference, amortized inference, and conditional generative modeling to handle supervised learning, Bayesian optimization, and simulation-based inference within a single framework. Experiments demonstrate that ACE significantly outperforms existing neural process methods on image completion and classification, black-box optimization, and simulation-based inference—achieving superior flexibility, probabilistic calibration, and cross-task generalization.

Explicit representation of interpretable latent variablesFlexible runtime conditioning on probabilistic latent informationImproved performance in optimization, simulation, and inference tasks

Controlled Training Data Generation with Diffusion Models

Mar 22, 2024
TY
Teresa Yeo
🏛️ Swiss Federal Institute of Technology Lausanne | MIT

This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.

Automate closed-loop feedback for adversarial prompt generationControl text-to-image models for supervised training dataGuide generation to match target data distributions

Latest Papers

What's happening recently
View more

This study addresses the persistent challenge in large language models (LLMs) of balancing controllability with fluency under conditional generation settings, a trade-off often overlooked by existing approaches that neglect output quality. The authors systematically evaluate representative techniques—including activation steering, prompt engineering, and supervised fine-tuning—on concept injection and removal tasks, employing both automated text metrics and LLM-as-a-judge assessments. Their analysis reveals that activation steering exhibits substantially degraded performance on instruction-tuned models, while highly effective control methods frequently compromise textual fluency. Furthermore, prompting and fine-tuning prove suitable for injecting concepts but struggle to reliably remove them. Notably, low-cost automatic metrics demonstrate strong correlation with human judgments, suggesting they can serve as viable, cost-efficient alternatives to expensive LLM-based evaluators.

concept injectionconcept removaleffectiveness-fluency trade-off

This work addresses the challenges in diffusion-based image editing, particularly the insufficient inversion accuracy and the trade-off between editing fidelity and background preservation. The authors propose SimEdit, a framework that investigates how textual conditions influence the geometry of the diffusion velocity field and cross-branch attention consistency, thereby revealing the critical role of condition precision in inversion stability. Building on this insight, they design a condition-aware dual-component editing mechanism that integrates condition refinement with token-wise cross-branch attention control. This approach substantially improves both image reconstruction quality after inversion and overall editing performance, outperforming existing attention-manipulation methods on the PIE-Bench benchmark.

background preservationdiffusion image editingediting fidelity

Existing activation-based intervention methods for text-to-image generation require per-concept optimization, limiting their applicability to open or dynamic concept sets. This work proposes HyperTransport, a framework that leverages a hypernetwork to directly map CLIP embeddings to intervention parameters, trained end-to-end with an optimal transport loss. HyperTransport enables single forward-pass generation of interventions for arbitrary novel concepts, unifying amortized intervention for open concept sets, continuously controllable interpretability strength, and cross-modal image-guided text generation for the first time. Experiments on DMD2 and Nitro-1-PixArt demonstrate that HyperTransport achieves generation quality comparable to per-concept optimization baselines across 167 unseen concepts, while accelerating inference by 3,600–7,000× and obtaining approximately twice the human and vision-language model (VLM) preference over prompt engineering.

activation steeringamortized controlconcept conditioning

This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.

controllable image generationdiffusion modelsgeneration control

This work addresses the absence of a unified classification framework for conditional 3D CT generation methods, which hinders systematic comparison and identification of critical design choices. To resolve this, the paper proposes a taxonomy centered on conditioning mechanisms, structured around three orthogonal dimensions: external knowledge type (Knowledge), knowledge integration paradigm (Integration), and generative architecture (Architecture), thereby defining a unified design space denoted as K×I×A. Through a comprehensive review of existing approaches, this framework not only clarifies prevailing technical pathways but also uncovers underexplored research directions. The resulting taxonomy offers theoretical guidance and a foundation for innovation in the design of future conditional 3D medical image generation methods.

3D CT generationconditional generationgenerative models

Hot Scholars

DC

Daniel C. Alexander

Professor of Imaging Science, Centre for Medical Image Computing, Department of Computer Science
Computer scienceMachine learningMedical imagingdiffusion MRI
JC

Jaemin Cho

PhD Student at UNC Chapel Hill
Multimodal LearningNatural Language ProcessingMachine Learning
MB

Mohit Bansal

Parker Distinguished Professor, Computer Science, UNC Chapel Hill
Natural Language ProcessingComputer VisionMachine LearningMultimodal AI
CC

Carlos Crispim-Junior

Associate Professor @ Université Lumière Lyon 2 - LIRIS UMR CNRS 5205
Artificial IntelligenceComputer VisionDeep LearningMultimodal Vision
ST

Sergey Tulyakov

Director of Research, Snap Inc.
computer visionmachine learning