IPVTON: Image-based 3D Virtual Try-on with Image Prompt Adapter

📅 2025-01-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address low garment-to-body geometric fidelity and texture distortion in image-driven 3D virtual try-on, this paper proposes a novel method integrating diffusion priors with explicit geometric constraints. Our approach features: (1) an image prompt adapter enabling cross-modal alignment between garment and human body from a single input image; (2) mask-guided image prompt embedding to emphasize semantic regions relevant to try-on; and (3) ControlNet-based pseudo-silhouette generation coupled with Score Distillation Sampling to optimize a hybrid SMPL+NeRF representation. Evaluated on standard benchmarks, our method achieves state-of-the-art performance—improving geometric accuracy by 18.7% (Chamfer Distance ↓) and texture realism by 22.3% (FID ↓). Qualitative results demonstrate substantial enhancements in garment wrinkling, draping behavior, and color consistency.

Technology Category

Computer Vision: Diffusion Models for VisionSearch and Optimization: Sampling/Simulation-based SearchHumans and AI: Game Design — Virtual Humans, NPCs and Autonomous Characters

Application Category

Economics, Online Markets and Human Computation: LLM based quality controls for crowd workUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Given a pair of images depicting a person and a garment separately, image-based 3D virtual try-on methods aim to reconstruct a 3D human model that realistically portrays the person wearing the desired garment. In this paper, we present IPVTON, a novel image-based 3D virtual try-on framework. IPVTON employs score distillation sampling with image prompts to optimize a hybrid 3D human representation, integrating target garment features into diffusion priors through an image prompt adapter. To avoid interference with non-target areas, we leverage mask-guided image prompt embeddings to focus the image features on the try-on regions. Moreover, we impose geometric constraints on the 3D model with a pseudo silhouette generated by ControlNet, ensuring that the clothed 3D human model retains the shape of the source identity while accurately wearing the target garments. Extensive qualitative and quantitative experiments demonstrate that IPVTON outperforms previous methods in image-based 3D virtual try-on tasks, excelling in both geometry and texture.
Problem

Research questions and friction points this paper is trying to address.

3D Virtual Fitting
Accuracy Improvement
Cloth Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

IPVTON
ControlNet
3D virtual try-on
🔎 Similar Papers
No similar papers found.