๐ค AI Summary
This study addresses the challenge of fine-grained, continuous control over style intensity in human motion diffusion models by proposing an endpoint-supervised motion style transfer framework. By integrating latent-space style direction construction, diffusion-based denoising, and latent intensity regularization, the method achieves smooth and monotonic style modulation using only scalar intensity values, eliminating the need for intermediate ground-truth annotations. Furthermore, the framework supports training on heterogeneous datasets and enables out-of-range extrapolation. Experimental results demonstrate that the proposed approach significantly enhances both the controllability and realism of generated motions while effectively preserving content consistency across interpolation and extrapolation tasks.
๐ Abstract
Existing human motion diffusion methods provide strong motion generation quality, and recent style transfer models can inject target style cues, but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit, but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised continuous control. Given a content motion and a style motion, we construct a style direction in a learned motion-style embedding space and condition diffusion generation with a scalar intensity. The training objective combines diffusion denoising with latent intensity regularization to encourage smooth and monotonic style scaling without requiring intermediate-intensity ground-truth motions. Our framework is compatible with pretrained motion diffusion backbones and supports heterogeneous style datasets, including the multi-actor style motion dataset. To test out-of-range usability, we additionally introduce a small real-capture over-reaction extension and evaluate large-intensity behavior against these unseen targets. Experiments measure controllability, interpolation/extrapolation behavior, content preservation, and motion realism, with ablations on direction construction and loss design.