DuoMatching: Joint-Marginal Distribution Matching for Few-Step Video Generation
This study addresses the challenges of insufficient visual quality and semantic misalignment in streaming video generation by proposing a distillation framework based on unified joint-marginal distribution matching. Methodologically, an image teacher model is leveraged to provide frame-level supervision for optimizing few-step generation. Furthermore, a LatentBridge mechanism is introduced to resolve cross-domain latent representation mismatches, combined with latent variant sampling to enhance dynamic representational capacity. Experimental results demonstrate that the proposed approach significantly improves both visual fidelity and text alignment while effectively preserving motion coherence, achieving a human preference rate exceeding 80%.