Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance

📅 2025-08-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In low-data regimes (e.g., CelebAMask-HQ), diffusion-based face generation suffers from weak semantic alignment and poor attribute controllability. To address this, we propose a multi-condition-guided controllable generation framework: (i) contrastive embedding learning via InfoNCE loss to enhance semantic consistency of attribute vectors; (ii) SegFormer as the segmentation encoder to strengthen fine-grained spatial-structural guidance; and (iii) a multimodal conditional input mechanism that fuses attribute vectors with segmentation masks, efficiently adapted to Stable Diffusion via LoRA fine-tuning. Experiments demonstrate significant improvements in both unconditional and conditional generation—achieving superior image fidelity and precise attribute control over mainstream baselines. Notably, our method exhibits enhanced generalization and controllability under data-scarce conditions.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: GenerationMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
We present a benchmark of diffusion models for human face generation on a small-scale CelebAMask-HQ dataset, evaluating both unconditional and conditional pipelines. Our study compares UNet and DiT architectures for unconditional generation and explores LoRA-based fine-tuning of pretrained Stable Diffusion models as a separate experiment. Building on the multi-conditioning approach of Giambi and Lisanti, which uses both attribute vectors and segmentation masks, our main contribution is the integration of an InfoNCE loss for attribute embedding and the adoption of a SegFormer-based segmentation encoder. These enhancements improve the semantic alignment and controllability of attribute-guided synthesis. Our results highlight the effectiveness of contrastive embedding learning and advanced segmentation encoding for controlled face generation in limited data settings.
Problem

Research questions and friction points this paper is trying to address.

Improving face generation with contrastive embeddings and segmentation guidance
Comparing UNet and DiT architectures for unconditional face generation
Enhancing controllability in limited-data face synthesis via multi-conditioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates InfoNCE loss for attribute embedding
Uses SegFormer-based segmentation encoder
Enhances semantic alignment and controllability
Dhruvraj Singh Rawat
Dhruvraj Singh Rawat
Grad Student, University of Surrey
E
Enggen Sherpa
University of Surrey, UK
R
Rishikesan Kirupanantha
University of Surrey, UK
T
Tin Hoang
University of Surrey, UK