Score
Design and implement temporal attention modules that use static pose or reference features as queries and motion-derived signals (micro-motion offsets, Doppler-energy, pose offsets) as keys/values, so attention weights emphasize dynamic changes over time. Build cross-attention temporal models that encode local relative velocities and motion offsets to suppress static body-shape biases and improve temporal alignment and motion sensitivity.
This work addresses the challenging problem of high-fidelity human motion generation conditioned on time-varying input signals—such as audio, text, or control commands. We propose Temporal-Conditional Mamba (TC-Mamba), the first method to introduce state-space models (SSMs) into conditional motion generation. Unlike prevailing approaches relying on cross-attention, TC-Mamba embeds conditioning information directly into the recurrent state updates of Mamba via selective state-space dynamics, enabling fine-grained, stepwise temporal alignment and motion control. By leveraging the SSM’s inherent long-range modeling capacity and linear-time inference, TC-Mamba significantly improves motion smoothness, temporal alignment accuracy, and condition fidelity—especially for extended sequences. It achieves state-of-the-art performance across multiple benchmark tasks, demonstrating both the effectiveness and superiority of state-space architectures for complex, temporally conditioned generative modeling.
本文提出COMET框架,通过显式时间表示、外观-运动融合及方向感知优化,解决了视频多模态大语言模型中细粒度运动-时间理解不足的问题。
This study addresses the challenge of decoupling temporal and control parameters during dynamic motion retargeting for legged robots with diverse morphologies. We propose a dense temporal motion retargeting method that introduces a novel dense temporal deformation mechanism. This mechanism jointly optimizes step-wise temporal and control parameters within a single optimization pass to precisely accommodate robot-specific dynamics, applying deformation only where necessary. Furthermore, the approach incorporates GPU-parallelized sampling-based model predictive control (MPC) to ensure computational efficiency. Experimental results demonstrate that our method achieves a 19-fold speedup over comparable baselines while yielding superior accuracy, and successfully transfers the retargeted policies to a real-world humanoid robot.
Existing text-driven 3D human motion generation methods often overlook the coupling between motion periodicity and keyframe saliency, and are sensitive to semantically equivalent textual rephrasings, leading to drift and instability in long-sequence generation. To address these limitations, this work proposes a periodicity- and saliency-aware Mamba architecture that integrates enhanced density peak clustering to estimate keyframe weights and FFT-accelerated autocorrelation for periodicity analysis. Furthermore, we introduce the Periodicity-Differential Cross-Modal Alignment Module (PDCAM), which explicitly models the coupling mechanism between periodicity and saliency for the first time, thereby enhancing the robustness of text-motion embedding alignment. Extensive experiments on HumanML3D and KIT-ML demonstrate state-of-the-art performance, achieving an FID of 0.068 and significantly outperforming existing approaches across all metrics.
This work addresses the challenge of balancing temporal consistency and detail preservation in Stable Diffusion–based video generation. The authors propose a motion-adaptive temporal attention mechanism that, without modifying the original frozen model, dynamically adjusts the attention receptive field based on motion estimation. This approach integrates a cascaded UNet injection strategy, temporally correlated noise initialization, and a motion-aware gating mechanism, introducing only 25.8 million (2.9%) additional trainable parameters. Experiments on the WebVid validation set demonstrate competitive generation quality, marking the first demonstration of high-fidelity video synthesis without explicit temporal loss terms. Furthermore, the method reveals a controllable trade-off between noise correlation and motion magnitude, offering new insights into motion-guided generative modeling.
Existing vision-language-action (VLA) models struggle to capture temporal dynamics due to their reliance on single-frame observations, leading to significant performance degradation in non-stationary environments. This work proposes a training-free, inference-time correction method that jointly and orthogonally adjusts rhythm compression and trajectory deviation for chunked-action VLA models via a closed-form solution, enabling low-latency and temporally consistent dynamic adaptation. Leveraging quadratic cost optimization and orthogonal decomposition, the approach achieves efficient dynamic compensation without any model retraining—a first in the field. Evaluated on the MoveBench benchmark, it improves task success rates by 28.8% in fully dynamic environments and by 25.9% in mixed static-dynamic settings, substantially outperforming existing training-free fine-tuning methods.
本文提出时间-频率几何交叉注意力(TFGCA)模块,解决视觉-语言-动作模型中动作序列的频率和跨阶段几何结构问题,提升模型性能。
Existing video Transformers struggle to model complete spatiotemporal dependencies in critical regions and long-range action dynamics due to their reliance on factorized or windowed self-attention mechanisms. Inspired by the human visual system’s “glance-and-gaze” strategy—characterized by an initial holistic glance followed by focused gaze—this work proposes the OG-ReG Transformer. The architecture incorporates a Glance pathway to capture global, coarse-grained spatiotemporal context and a Gaze pathway to attend to fine-grained local details, dynamically allocating sparse spatiotemporal attention and fusing multi-scale features. This approach introduces, for the first time, a dual-path visual attention mechanism into video understanding, moving beyond conventional uniform processing paradigms. It achieves state-of-the-art performance across multiple benchmarks, including Kinetics-400, Something-Something v2, and Diving-48.
This study addresses the limitation of existing video saliency benchmarks in distinguishing temporal models from static baselines, which leads to a severe underestimation of dynamic information. To this end, we propose SalTempto, a benchmark that curates highly dynamic, event-causality-focused clips from the HACS-Segments dataset and constructs an evaluation framework integrating multi-subject eye-tracking with pretrained model fine-tuning. Experimental results demonstrate that static models suffer significant performance degradation on this benchmark, whereas temporal architectures substantially improve prediction accuracy. Furthermore, our analysis reveals that nearly half of the potential performance margin remains unexplored. This work provides a critical benchmark for reassessing the true value of temporal modeling in video understanding.
为解决动态操作任务中的运动模糊和状态混淆问题,提出TEMPO方法,通过增加时间上下文信息来提高机器人动态操作的准确性。