DyMD: Preserving Interaction Dynamics through Distribution Matching Distillation in Few-Step Video World Models

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of interaction dynamics and degradation of motion quality in video diffusion model distillation by proposing the DyMD framework. This framework introduces temporally affinity-conditioned re-noise sampling and dynamics-guided pseudo-score tracking, integrated with adaptive teacher supervision, critic fitting, and latent temporal dynamics prediction to effectively balance motion recovery with appearance refinement. Experimental results demonstrate that the proposed method successfully distills a 14B-parameter model into a 1.3B-parameter model requiring only four inference steps. Furthermore, it preserves visual fidelity while improving task adherence by 9.6% and achieving a 34% success rate in action planning.
📝 Abstract
Large video diffusion models offer expressive priors for embodied prediction and learning, yet their many-step sampling remains costly for interactive downstream use. Distribution Matching Distillation (DMD) enables few-step video generation, but can suppress robot--object motion while preserving visual quality. Examining DMD's teacher and fake-score signals, we find that weak re-noising keeps the teacher posterior concentrated near motion-deficient rollouts, limiting motion-restoring guidance. Meanwhile, stronger-motion rollouts tend to incur larger fake-score fitting errors, which can hinder the generator's learning of interaction dynamics. We propose DyMD, a DMD framework that adapts both teacher supervision and critic fitting to the evolving student. Temporal affinity--conditioned re-noise sampling adapts the timestep distribution to each rollout's current interaction fidelity by mixing the base schedule with a teacher prior motivated by local posterior variation, thereby balancing motion recovery and appearance refinement. To better track stronger-motion rollouts, dynamics-guided fake-score tracking uses a noise-conditioned predictor to estimate noise-relative fitting difficulty from latent temporal dynamics, then upweights predicted-hard rollouts in the critic loss. Using DyMD, we distill a 14B teacher into a four-step 1.3B student with no auxiliary modules at inference. On embodied-video benchmarks, the student improves R-Bench task adherence by $9.6$ percentage points and PAI-Bench-G Domain score by $5.1$ points over Base DMD while maintaining comparable visual quality. As a backbone for downstream action planning, our student achieves 34% mean success across two WorldArena tasks, compared with 16% for Base DMD.
Problem

Research questions and friction points this paper is trying to address.

Video World Models
Distribution Matching Distillation
Interaction Dynamics
Embodied Prediction
Few-Step Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distribution Matching Distillation
Video World Models
Few-Step Generation
Interaction Dynamics
Embodied AI
🔎 Similar Papers
No similar papers found.