MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a high-fidelity, region-controllable makeup transfer method that accurately preserves the identity of the source face while effectively transferring target makeup styles. Built upon a diffusion Transformer architecture, the approach introduces a novel Token-Aligned Region Gating mechanism to enable pixel-level and attention-aligned regional control. Furthermore, Cross-Modal Perception Guidance is employed to effectively disentangle makeup and identity representations. The study also presents the first high-resolution makeup dataset with region-level annotations and establishes a unified evaluation benchmark. Experimental results demonstrate that the proposed method significantly outperforms existing approaches in terms of regional controllability, makeup fidelity, and identity preservation, exhibiting strong robustness across diverse makeup styles, ethnicities, and facial poses.
📝 Abstract
Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are (i) pixel-to-attention misalignment that causes spillover into non-target areas and weakens regional control; (ii) unclear transfer/preservation concept separation under two-image conditioning, leading to coupling between makeup attributes and identity; and (iii) the lack of a high-resolution dataset that is identity-consistent and region-labeled for fine-grained supervision. In this paper, we propose MagicMakeup, a diffusion transformer-based framework for region-controllable and high-fidelity makeup transfer, built on spatial constraints and concept disentanglement. To enable precise region-specific editing while preserving identity, we propose Token-Aligned Region Gating, which aligns pixel masks with attention and applies region-specific logit gating. To clarify the concepts of transfer and preservation, we further introduce Cross-Modal Perception Guidance, which aligns text and image features to enhance cross-modal concept perception. We also design a pipeline for the generation of 1024 x 1024 data pairs through region-specific makeup removal and establish a unified benchmark in synthetic and real settings. Extensive quantitative and qualitative experiments show that MagicMakeup improves regional controllability, makeup fidelity, and identity preservation, with strong robustness across styles, races, and poses.
Problem

Research questions and friction points this paper is trying to address.

makeup-transfer
regional controllability
identity preservation
diffusion models
concept disentanglement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Region-Controllable Diffusion
Token-Aligned Region Gating
Cross-Modal Perception Guidance
Concept Disentanglement
High-Fidelity Makeup Transfer
🔎 Similar Papers