π€ AI Summary
This study addresses the detail degradation in existing high-resolution video virtual try-on methods, which stems from insufficient garment information utilization and the absence of cross-modal positional modeling. To overcome these limitations, this work proposes a high-fidelity try-on framework built upon a pretrained video diffusion Transformer. The core innovations include a timestep-adaptive modulation mechanism, a frame-aligned positional encoding strategy, and a multi-source condition injection design, collectively enhancing local correspondence while suppressing conditional interference. Experimental evaluations demonstrate that the proposed method achieves state-of-the-art performance on benchmarks such as Eevee, excelling in garment detail preservation, temporal consistency, and overall video quality.
π Abstract
Video virtual try-on has attracted increasing attention due to its broad potential in digital fashion and intelligent e-commerce. However, existing methods primarily focus on low-resolution settings and still face substantial challenges when extended to high-resolution scenarios. These limitations can be attributed to two main factors: (1) the insufficient utilization of rich garment reference information, and (2) the lack of explicit positional modeling between garment and video representations during cross-modal interaction, which weakens fine-grained local correspondence. To address these issues, we propose TexTailor, a high-fidelity video virtual try-on framework built upon a pretrained video Diffusion Transformer. Specifically, we introduce a timestep-adaptive modulation mechanism to dynamically adjust garment visual representations throughout denoising. We further develop a frame-aligned positional encoding strategy to strengthen garment-to-video correspondence, together with a multi-source injection design that reduces interference among heterogeneous conditions. Extensive experiments on multiple video virtual try-on benchmarks, including the high-resolution Eevee dataset, demonstrate that TexTailor achieves competitive performance in garment detail preservation, temporal consistency, and overall video quality.