DynaConTalk: Wavelet-Constrained Diffusion for Long-Form and Controllable Holistic Co-Speech 3D Motion

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-smoothing and detail loss caused by motion representation limitations and speech-conditioned averaging in co-speech animation. To overcome these challenges, this work proposes a wavelet-constrained diffusion framework that pioneers the decoupling of multi-band coarse-to-fine motions within the stationary wavelet domain. The method extracts multimodal features using HuBERT with attention pooling and achieves adaptive fusion through a noise-aware dynamic gating network. Furthermore, it introduces a frame-level rhythm pathway to support long-horizon controllable generation, historical continuation, and key-pose inpainting. Experimental results demonstrate that the proposed framework significantly improves facial accuracy and temporal continuity while providing an interactive editing interface. The code and pretrained models have been made publicly available.
📝 Abstract
Holistic co-speech animation is prone to averaging in both motion representation and speech conditioning. In coordinate-space diffusion, slow body posture, mid-frequency gesture strokes, and fast hand or facial details are entangled in one prediction target, often producing low-variance, over-smoothed motion. Meanwhile, dense rhythmic and acoustic cues can dominate sparse content-specific information under fixed multimodal fusion. We present DynaConTalk, a wavelet-constrained diffusion framework for long-form and controllable holistic co-speech motion generation. Diffusion operates in stationary wavelet transform (SWT) coefficient space, whose temporally aligned bands separate coarse posture evolution, gesture strokes, and fine expressive details. Our dynamic gating network preserves HuBERT and speaker identity as a base and selectively adds rhythm, mel, and transcript features through motion-state- and noise-aware residual gates. Attention pooling and learned depth routing deliver complementary conditions to each denoising stage, while a frame-resolution rhythm path preserves precise timing. A signed proposal-consensus update then reconciles these conditions with the evolving motion state. Matched-noise constraint injection uses the same sampling interface for history continuation and localized keypose repair, and extends to reference-guided control. Separate body-hand and facial denoisers, followed by inverse SWT and a pose-driven root regressor, produce holistic motion. Experiments evaluate generation quality, facial accuracy, temporal continuity, and controllable editing. Code, models, and the interactive editing interface are available at https://github.com/zhuyifeiabcd1/DynaConTalk.
Problem

Research questions and friction points this paper is trying to address.

co-speech 3D motion generation
over-smoothed motion
multimodal fusion
holistic animation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Wavelet-Constrained Diffusion
Co-Speech 3D Motion
Dynamic Gating Network
Matched-Noise Constraint Injection
Holistic Animation
🔎 Similar Papers
No similar papers found.