LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

📅 2026-08-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决视频虚拟试穿的实时性问题,提出LiveVVT框架,通过滚动流式扩散和双记忆机制保持高保真度,同时显著降低延迟。
📝 Abstract
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip dependence incurs prohibitive latency and computational overhead in practical continuous deployment. Naively enforcing causality disrupts pretrained bidirectional priors and substantially degrades synthesis quality. We introduce LiveVVT, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation. Within a fixed-size window, LiveVVT jointly denoises multiple video chunks under bounded look-ahead, preserving local bidirectional interactions while emitting one clean chunk per iteration. Beyond the window, two complementary memories sustain long-term consistency: a bounded temporal memory propagates recent dynamics and occlusion context, whereas a persistent global appearance memory, constructed once from the target garment and a frontal try-on keyframe, anchors garment details and dressed appearance throughout the stream. We further introduce a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference. Experiments on paired and unpaired long-sequence benchmarks demonstrate superior generation quality over similarly sized models, with $26\times$ lower latency and $11\times$ higher throughput, enabling high-fidelity real-time streaming VVT.
Problem

Research questions and friction points this paper is trying to address.

Video Virtual Try-On
diffusion-based
causality
bidirectional priors
latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

rolling streaming diffusion
bounded bidirectional modeling
causal recurrent generation
complementary memories
progressive distillation framework
Y
Yushe Cao
Tsinghua University
Shikun Feng
Shikun Feng
Baidu
nlp
R
Ruxiang Duan
South China University of Technology
L
Liyong Wang
Beijing Jiaotong University
D
Dianxi Shi
Tsinghua University
C
Chun Yu
Tsinghua University
J
Junliang Xing
Tsinghua University