Triangular Resampling for Long-Horizon Motion Generation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the error accumulation in long-horizon motion generation with diffusion models, caused by mismatches between training and inference states. To mitigate this, we propose a triangular resampling framework that extends roll-out training to partially denoised states while unifying denoising thresholds. By integrating a FloodDiffusion-based triangular schedule, gradient-free multi-step replay, and a noise-matched ground-truth clamping mechanism, our approach effectively bridges the discrepancy between training windows and generative states. Evaluated on HumanML3D, the proposed method achieves state-of-the-art FID AUC performance. Notably, supervised triangular resampling reduces FID AUC by 40.9% and the degradation slope by 55.3%, substantially improving the stability of long-term motion generation.
📝 Abstract
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference states. Replacing only completed motion history leaves this mismatch unresolved in partially denoised states within the active window. TR therefore extends rollout-based training to these states, using ground-truth clamping to limit excessive drift. For each replayed sample, TR draws one denoising threshold, shared across latent positions and replay updates, and replays multi-step triangular denoising without gradient tracking. After each update, states below the threshold are replaced with noise-matched ground truth, while those at or above it retain model predictions. The resulting latent window enters the standard training update. This rollout construction supports both supervised training (TR) and distribution matching (TR-DMD). On 120-second motion generation from HumanML3D test prompts, TR and TR-DMD achieve state-of-the-art FID AUC within their respective non-DMD and DMD comparison groups. Supervised TR reduces FID AUC by 40.9% and FID degradation slope by 55.3% relative to matched post-training without replay.
Problem

Research questions and friction points this paper is trying to address.

long-horizon motion generation
error accumulation
motion diffusion models
training-inference mismatch
Innovation

Methods, ideas, or system contributions that make the work stand out.

Triangular Resampling
Motion Diffusion Models
Long-Horizon Motion Generation
Error Accumulation Mitigation
Distribution Matching