Score
Designs and implements probabilistic models that represent distributions over agent motion and future trajectories, producing predicted position heatmaps, sampled trajectory realizations, or per-subject motion likelihoods. Builds and evaluates generative motion models that capture within-subject variability and can be conditioned on or used to assess identity-related motion for purposes such as prediction, projection, and identity inference.
The field of motion generation lacks a systematic survey grounded in generative methodology. To address this gap, we propose the first deep taxonomy centered on generative strategies, synthesizing state-of-the-art works from top-tier conferences (CVPR, ICCV, SIGGRAPH, CoRL) since 2023. Our framework uniformly analyzes four dominant paradigms—GANs, autoencoders, autoregressive models, and diffusion models—across three dimensions: architectural design, conditional modeling mechanisms, and evaluation protocols for motion sequence synthesis. We consolidate widely adopted datasets and metrics, identifying key challenges including motion coherence, physical plausibility, and cross-domain generalization. By establishing a comprehensive, comparable analytical benchmark with well-defined dimensions, our survey significantly enhances methodological comparability and facilitates precise problem diagnosis. This work serves as a foundational reference for researchers advancing generative motion modeling.
This study addresses the limited effectiveness of existing methods in identity recognition based on individual motion styles, particularly in surveillance and authentication scenarios. The work proposes a novel interactive identity recognition paradigm that infers identity through sequences of visual prompts from the system and corresponding motion responses from the subject. Grounded in human information processing mechanisms, the approach employs a probabilistic generative model combined with Bayesian posterior updating and mutual information optimization to dynamically select the most discriminative visual stimuli, thereby maximizing mutual information between prompts and identity. Evaluated across five public datasets and a newly collected dataset comprising 4,476 records, the method achieves consistently high accuracy, significantly enhancing both recognition efficiency and robustness.
In human-robot collaboration scenarios (e.g., rehabilitation, sports, manufacturing), virtual avatars and robots require realistic, individualized motion replication—yet existing models fail to capture subject-specific kinematic traits such as velocity distributions and amplitude envelopes. Method: We propose the first fully data-driven framework for personalized human motion generation, employing an LSTM-based temporal generative model trained end-to-end on scalar oscillatory motion data from real human subjects. Contribution/Results: Our approach accurately encodes individual rhythmic and amplitude characteristics while preserving inter-subject variability. Quantitative evaluation demonstrates superior motion similarity over current state-of-the-art methods. Crucially, it achieves the first distinguishable and reproducible modeling of subject-specific movement patterns—establishing a foundational motion modeling capability for natural group-level interactions in XR avatars and embodied agents.
This paper addresses the problem of single-image motion trajectory forecasting: predicting dense future trajectories of scene objects directly from a static image, without requiring auxiliary physical parameters such as velocity or force. We propose a conditional generative model based on a *trajectory grid*, which bypasses redundant pixel-level modeling typical in video generation and instead performs end-to-end synthesis of structured motion fields. Our approach explicitly captures global dynamic patterns and motion uncertainty. Built upon modern video generation architectures, the model is trained jointly on synthetic physics-based simulations and real-world scenes. Experimental results demonstrate significant improvements over state-of-the-art regression- and generation-based methods on both simulated and real-world physical benchmarks. Furthermore, we validate the practical utility and generalization capability of our method in downstream robotic navigation tasks.
This study addresses fundamental challenges in human motion generation—specifically, motion representation design and loss function formulation. We propose vMDM, a lightweight surrogate diffusion model, to systematically evaluate six mainstream motion representations across multiple datasets in terms of generation quality, diversity, and training efficiency. Crucially, we introduce v-loss—a unified prediction objective—for the first time. Through controlled ablation experiments, we reveal that motion representation choice critically governs latent-space distribution modeling capability and conditional generation performance. Notably, representations combining joint velocities with rotation matrices yield substantial improvements: up to 2.3× faster convergence and significantly lower FID scores. Our findings provide both theoretical insights and empirical guidelines for motion representation selection and loss optimization in diffusion-based motion generation frameworks.
Existing approaches struggle to efficiently model the multimodal distribution of future scene dynamics from partial observations. To address this challenge, this work proposes GARFIELD, a novel framework that explicitly constructs a structured spatiotemporal latent variable distribution to represent all plausible future motions conditioned on an input image and sparse constraints. By integrating object-aware latent space modeling, a deterministic density decoder, and probabilistic trajectory generation, GARFIELD enables joint sampling, local uncertainty estimation, progressive constraint fusion, and interactive density queries—without requiring large-scale video generation or Monte Carlo sampling. Experiments demonstrate that GARFIELD achieves trajectory sampling 97× faster than large video generation models and density estimation two orders of magnitude faster, while delivering competitive performance in motion planning tasks and supporting real-time interaction.
This work proposes a probabilistic inference framework that integrates inductive biases to address the challenges of uncertainty quantification in deep sequential models. While traditional Bayesian approaches struggle with prior specification and inference accuracy in large-scale networks, the proposed method establishes a theoretical connection between Transformer attention mechanisms and sparse Gaussian processes, enabling scalable approximate Bayesian inference. It introduces cross-domain inducing points derived from HiPPO operators to support long-range historical modeling in online learning settings. Furthermore, self-supervised signals are leveraged to enrich the probabilistic structure of latent variables in sequence generation. The resulting approach significantly enhances the uncertainty quantification capability, probabilistic expressiveness, and scalability of deep sequential models, all while maintaining competitive predictive performance.
This study addresses the lack of systematic evaluation of uncertainty reliability in probabilistic forecasting for physical systems, particularly between generative models and ensemble methods trained with the Continuous Ranked Probability Score (CRPS). The authors establish a unified evaluation framework to conduct the first systematic comparison of these two approaches under identical model scales and computational budgets in two-dimensional spatiotemporal systems. Results demonstrate that CRPS-based ensembles consistently achieve more reliable uncertainty coverage and faster inference in both single-step and roll-out predictions. Generative models exhibit comparable coverage only when trained in the original data space but incur significantly higher latency. Both approaches yield similar predictive accuracy. Notably, CRPS ensembles maintain strong performance even in latent spaces, highlighting their advantage in enabling efficient and reliable uncertainty quantification.
This work addresses the limitations of existing methods that overly rely on dense appearance representations, which struggle to efficiently model long-horizon, multimodal sparse motion trajectories in complex scenes. The authors propose a dynamics-centric sparse point trajectory representation and formulate its evolution as an autoregressive diffusion process, enabling stepwise local predictions that explicitly capture the temporal accumulation of uncertainty. Their approach generates thousands of diverse, physically plausible future trajectories from a single input image and introduces a new benchmark, OWM, for evaluating predictive distributions in open-world settings. Experiments demonstrate that the method matches or exceeds the prediction accuracy of current dense simulators while achieving sampling speeds several orders of magnitude faster, thereby enabling scalable and practical open-set future scene prediction.