🤖 AI Summary
This study addresses the entanglement of multi-scale motion frequencies and detail loss caused by fixed temporal grids by proposing the FreqMo framework. This method leverages wavelet decomposition to decouple motion into distinct frequency bands, enabling adaptive scale representation. Furthermore, it introduces a novel Unified Frequency Residual Quantization (UFRQ) scheme that encodes the full frequency spectrum within a single codebook, achieving three-fold compression of token sequences and stable single-stage generation. Integrated with diffusion models, the proposed approach attains state-of-the-art fidelity, significantly enhances high-frequency detail preservation, and seamlessly transfers to continuous diffusion backbones.
📝 Abstract
Most human motion generation methods encode motion as tokens on a uniform temporal grid, where every token spans the same fixed time window. Human motion, however, is temporally heterogeneous: slowly evolving global trajectories coexist with rapid transient events such as foot contacts and joint impulses. Forcing such multi-scale dynamics onto tokens of identical temporal resolution entangles motion frequencies, leaving slow regions redundant while smoothing out the rapid details that distinguish realistic motion. We propose \textbf{FreqMo}, a scale-adaptive motion representation that decomposes motion into wavelet frequency bands, separating dynamics across temporal scales while preserving temporal localization and exact reconstruction. Unified Frequency Residual Quantization (UFRQ) then encodes all bands within a single shared codebook, compressing the token sequence threefold and enabling stable single-stage generation. Experiments show FreqMo attains SOTA fidelity with substantially improved high-frequency preservation, and the same decomposition transfers to continuous diffusion backbones.