SuperMotion: Source-Preserving Denoising for Text-Driven Human Motion Editing

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of temporal details in unedited regions during text-driven motion editing with diffusion models, where reliance on conditional learning causes progressive detail loss throughout denoising. To this end, we propose a source-preserving denoising framework that employs a learnable gating mechanism to predict retention ratios and explicitly fuses authentic source motions as anchors during reverse steps, thereby injecting sample-level temporal details into the sampling posterior to correct estimations. Furthermore, a temporal high-frequency loss is introduced for supervision, effectively preserving dynamic features without requiring explicit editing masks. Evaluated on the MotionFix dataset, the proposed method achieves an R@1 editing accuracy of 33.20%, significantly mitigating temporal detail loss while enhancing the preservation of motion dynamics.
📝 Abstract
Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20\% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.
Problem

Research questions and friction points this paper is trying to address.

Human Motion Editing
Diffusion Models
Temporal Detail Preservation
Source-Preserving Denoising
Innovation

Methods, ideas, or system contributions that make the work stand out.

Source-Preserving Denoising
Preservation Gate
Clean-Space Source Anchor
Text-Driven Motion Editing
Temporal High-Frequency Loss
🔎 Similar Papers
No similar papers found.