Learning Structured Output Representations from Attributes using Deep Conditional Generative Models

📅 2023-04-30
🏛️ arXiv.org
📈 Citations: 8
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses three key challenges in structured image generation: (1) imprecise attribute control, (2) blurry outputs, and (3) oversimplified, unimodal prior modeling. To this end, we propose a multimodal disentangled generative framework based on conditional variational autoencoders (CVAEs). Methodologically, we introduce a weighted evidence lower bound (ELBO) optimization strategy that explicitly models multimodal priors over fine-grained attributes—such as hair color, eyewear presence, and species-specific traits—and enforces attribute-disentangled representations in the latent space. Notably, this is the first systematic application of CVAEs to attribute-driven, cross-domain structured generation (spanning faces and birds). Experiments on CelebA and CUB-200-2011 demonstrate substantial improvements in attribute accuracy, sample diversity, visual fidelity, and cross-category generalization, while maintaining robustness and interpretability.
📝 Abstract
Structured output representation is a generative task explored in computer vision that often times requires the mapping of low dimensional features to high dimensional structured outputs. Losses in complex spatial information in deterministic approaches such as Convolutional Neural Networks (CNN) lead to uncertainties and ambiguous structures within a single output representation. A probabilistic approach through deep Conditional Generative Models (CGM) is presented by Sohn et al. in which a particular model known as the Conditional Variational Auto-encoder (CVAE) is introduced and explored. While the original paper focuses on the task of image segmentation, this paper adopts the CVAE framework for the task of controlled output representation through attributes. This approach allows us to learn a disentangled multimodal prior distribution, resulting in more controlled and robust approach to sample generation. In this work we recreate the CVAE architecture and train it on images conditioned on various attributes obtained from two image datasets; the Large-scale CelebFaces Attributes (CelebA) dataset and the Caltech-UCSD Birds (CUB-200-2011) dataset. We attempt to generate new faces with distinct attributes such as hair color and glasses, as well as different bird species samples with various attributes. We further introduce strategies for improving generalized sample generation by applying a weighted term to the variational lower bound.
Problem

Research questions and friction points this paper is trying to address.

Mapping low to high dimensional outputs
Reducing spatial information loss
Controlled generation using attributes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Variational Auto-encoder framework
Disentangled multimodal prior distribution
Weighted term in variational lower bound
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
McGill University
M
Mohamed Debbagh
McGill University, Montreal, Quebec