🤖 AI Summary
In low-data regimes (e.g., CelebAMask-HQ), diffusion-based face generation suffers from weak semantic alignment and poor attribute controllability. To address this, we propose a multi-condition-guided controllable generation framework: (i) contrastive embedding learning via InfoNCE loss to enhance semantic consistency of attribute vectors; (ii) SegFormer as the segmentation encoder to strengthen fine-grained spatial-structural guidance; and (iii) a multimodal conditional input mechanism that fuses attribute vectors with segmentation masks, efficiently adapted to Stable Diffusion via LoRA fine-tuning. Experiments demonstrate significant improvements in both unconditional and conditional generation—achieving superior image fidelity and precise attribute control over mainstream baselines. Notably, our method exhibits enhanced generalization and controllability under data-scarce conditions.
📝 Abstract
We present a benchmark of diffusion models for human face generation on a small-scale CelebAMask-HQ dataset, evaluating both unconditional and conditional pipelines. Our study compares UNet and DiT architectures for unconditional generation and explores LoRA-based fine-tuning of pretrained Stable Diffusion models as a separate experiment. Building on the multi-conditioning approach of Giambi and Lisanti, which uses both attribute vectors and segmentation masks, our main contribution is the integration of an InfoNCE loss for attribute embedding and the adoption of a SegFormer-based segmentation encoder. These enhancements improve the semantic alignment and controllability of attribute-guided synthesis. Our results highlight the effectiveness of contrastive embedding learning and advanced segmentation encoding for controlled face generation in limited data settings.