Score
Designs and implements shared latent embedding spaces that encode motion-related signals—such as pose sequences, camera trajectories, video features, and text—and learns encoders and cross-modal mappings so these modalities occupy a common motion latent space. Builds models and tools to predict and manipulate motion codes (including topology-agnostic latents), to condition generative models via conditional flow matching or video-to-animation flow matching, and to support cross-modal interpolation, composition, and retargeting without test-time optimization.
This work addresses the limited generalization of trajectory prediction models in cross-dataset scenarios, primarily caused by discrepancies in scene layouts, agent behaviors, and perceptual conditions. The authors propose a transferability assessment framework based on latent scene embeddings and distributional distance metrics, establishing the first large-scale transfer experiment suite encompassing 24 mainstream trajectory datasets. Their analysis reveals a strong correlation between inter-dataset similarity and model transfer performance, enabling the design of a transferability scoring metric that effectively predicts cross-domain model behavior. This metric provides both theoretical grounding and practical guidance for pretraining strategies, dataset selection, and the development of foundational models for trajectory prediction.
This work addresses the computational inefficiency of conventional video synthesis methods in long-horizon motion generation by proposing an efficient generative framework based on highly compressed motion embeddings. The approach learns an implicit motion representation from large-scale trajectory data, achieving a temporal compression ratio of 64×, and constructs a conditional flow-matching model within this compressed latent space to flexibly respond to text prompts or spatial perturbations. Notably, it is the first method to directly model dynamics in the compressed embedding space, substantially improving both generation efficiency and controllability. Experiments demonstrate that the generated motions surpass state-of-the-art video generation models and specialized motion synthesis approaches in terms of realism, diversity, and computational efficiency.
This work addresses the challenge of flexibly supporting multimodal camera motion control in video generation. To this end, the authors propose a modality-agnostic framework that maps video, pose, and text inputs into a unified motion embedding space, enabling consistent and precise viewpoint manipulation. Key contributions include the construction of a Motion Triplet Dataset, the introduction of a geometry-driven motion representation based on camera extrinsics, and the design of a motion consistency objective in the latent space. The proposed method not only unifies multimodal inputs under a single processing pipeline but also enables novel capabilities such as motion sequence composition and cross-modal interpolation. Experiments demonstrate that the approach generates high-quality videos across all three modalities, accurately adhering to target camera trajectories and validating its effectiveness and generalization.
Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.
Existing methods struggle to generalize motion representations across species and characters with significantly different skeletal topologies, hindering the development of scalable generative models. This work proposes a semantic-aware, topology-agnostic motion representation framework that decouples motion from skeletal structure by aligning functionally corresponding joints through a semantic modulation mechanism, thereby constructing a unified latent motion manifold. The approach enables, for the first time, zero-shot cross-species motion retargeting without requiring paired data and supports learning a continuous, generation-friendly motion space directly from large-scale, unaligned raw BVH sequences. Experiments demonstrate high-fidelity motion reconstruction on both human and animal datasets, with successful applications in text-to-motion generation and cross-species motion transfer.
This work addresses the instability and motion distortion commonly observed in text-to-motion generation when directly employing latent representations from self-supervised encoders such as Motion-JEPA, which often fail to simultaneously preserve semantic accuracy, temporal coherence, and physical plausibility. To overcome these limitations, the authors propose MoRAE, a novel framework that systematically identifies and mitigates two key bottlenecks: ill-conditioned spectral properties in the latent space and misalignment between latent geometry and decoder-sensitive directions. MoRAE introduces compact bottleneck distillation to refine the spectral structure of JEPA features and incorporates motion-coupled training to align the latent manifold with decoder sensitivities. Integrated with Flow-Matching and a DiT architecture, the method achieves state-of-the-art performance under standard non-autoregressive settings, significantly enhancing both the quality and stability of generated motions.
This work investigates which pretraining signals induce action-relevant structural information in the latent spaces of video world models. Employing a unified inverse dynamics probing framework, the study systematically evaluates diverse pretraining strategies—including image-based self-supervision, video temporal modeling, autoencoders, diffusion models, and explicit dynamics models. The findings demonstrate that pretraining leveraging natural video temporal context substantially outperforms approaches focused on high-fidelity pixel reconstruction, revealing that temporal predictive structure—not reconstruction fidelity—is key to learning action-grounded visual representations. The resulting models exhibit superior generalization and robustness to visual perturbations across multiple robotic benchmarks, achieving a more favorable trade-off between visual fidelity and action prediction capability.
Existing motion capture methods are limited by fixed skeletal templates or reliance on cumbersome manual rigging, hindering generalization to characters with arbitrary topologies. This work proposes TopoCap, a unified framework that, for the first time, enables motion extraction from monocular video and zero-shot retargeting to any unknown skeletal topology—including bipeds, hexapods, and even inanimate objects—without test-time optimization. The approach leverages a graph-conditioned variational autoencoder to learn a universal motion prior and combines structural embeddings with conditional flow matching to map visual inputs to topology-agnostic motion codes. TopoCap outperforms specialized models on both human and quadruped benchmarks and successfully animates long-tail 3D characters in a zero-shot setting. The study also introduces Mobjaverse, a large-scale dataset encompassing over 5,000 distinct topologies and 2 million motion frames.
Existing motion capture datasets suffer from limited diversity, which constrains the generalization capabilities of generative models on rare, highly dynamic, and compositionally complex actions. To address this limitation, this work proposes a method that leverages large-scale synthetic human motion data combined with physics-based plausibility constraints to jointly expand both the training distribution and the size of the discrete codebook. By reconstructing the VQ-VAE motion tokenizer beyond the confines of real-data distributions, the approach substantially broadens the coverage and compositional capacity of the discrete motion representation space. This leads to consistent performance gains in text-to-motion generation and motion in-betweening tasks, and the enhanced representations can be seamlessly integrated into existing frameworks such as MotionGPT, demonstrating both the effectiveness and generalizability of the proposed representation expansion.