TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the deployment inefficiency of text-to-motion generation models caused by multi-step denoising and large parameter counts. We propose Terminal-Amplified Control Distillation (TACD), which trains an efficient student model without real data via piecewise online policy flow distillation. Furthermore, TACD introduces a terminal-amplified control mechanism that binds teacher queries to student step sizes to constrain loss weights, effectively correcting the error-weighting failure mode of fixed supervision grids at the denoising endpoint. Experiments demonstrate that TACD achieves 7.7–11.9× end-to-end speedup and 3.8–6.7× VRAM reduction, with significantly improved FID scores and retrieval performance matching or surpassing that of the teacher model.
📝 Abstract
Recent text-to-motion models have improved motion quality and instruction following, yet many-step denoising and large model components make deployment slow and memory-intensive. We present Terminal-Amplification-Controlled Distillation (TACD), an on-policy approach for training efficient motion generators from text prompts and pretrained teachers, without real-motion training data. Building on segmented on-policy flow distillation, we supervise clean-motion predictions along student-generated trajectories. We identify a failure mode in which velocity matching on a fixed supervision grid repeatedly overweights errors near the denoising endpoint, degrading few-step generation. TACD ties the latest teacher query to the student's step size, bounding the effective loss weights in clean-motion space without changing inference. Experiments on HumanML3D and KIT-ML demonstrate improved few-step generation, including a 58% reduction in eight-step HY-Motion student FID relative to distillation without this bound. For diffusion teachers, the endpoint-matching form of TACD yields four-step students with lower FID and matched or improved text-motion retrieval relative to their 50-step teachers on HumanML3D. On HY-Motion and Kimodo, eight-step students with compact components achieve 7.7-11.9x end-to-end speedups and reduce peak GPU memory by 3.8-6.7x relative to their teachers. Project page: https://vkgo.github.io/TACD/
Problem

Research questions and friction points this paper is trying to address.

text-to-motion
model distillation
few-step generation
deployment efficiency
denoising endpoint error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Terminal-Amplification-Controlled Distillation
Text-to-Motion
Flow Distillation
Few-step Generation
On-policy Distillation
🔎 Similar Papers
No similar papers found.