Parasitic Co-Denoising: Unlocking 3D Human Motion Generation in a Frozen Video Diffusion Model

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses how to distill the implicit knowledge of frozen video diffusion models into explicit 3D human motion without training separate auxiliary models. To this end, it proposes a parasitic co-denoising paradigm that employs Wan2.1 as the host model and reuses its noise schedule. By introducing a lightweight parasitic motion decoder alongside a Οƒ-adaptive multi-layer feature fusion mechanism, motion signals are extracted directly from the host’s intermediate representations while keeping all host parameters frozen. This work demonstrates that minimal trainable parameters suffice to unlock dispersed implicit motion priors within pretrained models, enabling the simultaneous generation of high-quality paired videos and 3D motions in a single inference pass. The proposed approach achieves text-to-motion alignment superior to dedicated generators while significantly outperforming existing baselines in efficiency.
πŸ“ Abstract
Despite never being supervised on explicit 3D motion, large-scale text-to-video diffusion models synthesize realistic human motion in their generated videos. We ask whether this implicit knowledge can be turned into explicit 3D motion generation, without training a separate motion model. Probing a frozen Wan2.1 reveals that a recoverable motion signal is present in its intermediate states across the entire denoising schedule, not confined to the clean output. Motivated by this, we introduce parasitic co-denoising, a paradigm in which motion is decoded from the host model along its denoising schedule rather than produced by an independent generator. We instantiate it as the Parasitic Motion Decoder (PMD), an efficient flow-matching decoder that shares the host's noise schedule and reads its intermediate features through a $Οƒ$-adaptive multi-layer fusion, leaving the host unmodified. Drawing its coverage from the host rather than from motion data, PMD leads dedicated motion generators on text-motion alignment at a small fraction of their trainable parameters, while producing paired video and motion in a single pass that motion-only baselines cannot match.
Problem

Research questions and friction points this paper is trying to address.

3D Human Motion Generation
Text-to-Video Diffusion Models
Implicit Knowledge Extraction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parasitic Co-Denoising
3D Human Motion Generation
Frozen Video Diffusion Model
Flow-Matching Decoder
Sigma-Adaptive Multi-Layer Fusion