Score
Designs and trains probabilistic priors over motion and action sequences—multimodal, history‑conditioned, latent‑variable, and population-level—so they can be sampled, conditioned, or decoded into continuous kinematic trajectories or implicit neural representations for bodies, hands, or other articulated systems. This includes specifying loss formulations and estimators (e.g., IMLE/implicit MLE), enforcing kinematic and contact consistency, conditioning on recent history or semantic cues, and mapping latent codes to time-indexed pose or motion outputs.
This work addresses the challenge of modeling probabilistic distributions of human poses represented as 6D rotations—i.e., elements on the SO(3) manifold. We propose a normalizing flow-based neural prior that explicitly respects the geometric structure of SO(3). Methodologically, we introduce the first integration of the RealNVP architecture with invertible Gram–Schmidt orthogonalization, enabling manifold-aware density estimation for 6D rotation representations. Our novel inverse Gram–Schmidt procedure ensures strict satisfaction of rotational constraints while preserving bijectivity and tractable likelihood evaluation. The resulting model exhibits strong representational capacity, training stability, and framework-agnostic compatibility. Experiments demonstrate substantial improvements in pose prior accuracy and cross-scenario generalization on motion capture and reconstruction tasks. Our approach establishes a new paradigm for probabilistic human motion modeling grounded in differential geometry and deep generative learning.
This study investigates how humans efficiently recognize everyday activities from complex visual inputs and disentangles the critical motion information required for classification versus action reconstruction. By comparing three representation methods—Temporal Movement Primitives (TMP), Legendre polynomial coefficients, and autoencoders—the contributions of static body poses and temporal dynamics are analyzed across videos of 16 daily activities. The findings reveal that static poses alone suffice for high-accuracy activity classification, with nine key joints identified as most discriminative, whereas naturalistic action reconstruction fundamentally relies on temporal dynamics. TMP and Legendre coefficients achieve comparable, top performance in classification; however, only TMP generates perceptually natural motions, highlighting an essential decoupling between the informational demands of action recognition and action generation.
To address poor generalization, motion jitter, and difficulty in sharing priors in physics-based character motion control, this paper proposes a continuous-discrete hybrid latent representation architecture: continuous residual modeling captures fine-grained motion details, while Residual Vector Quantization (RVQ) enhances the expressiveness and diversity of discrete latent variables, effectively mitigating posterior collapse. The method integrates VAE extensions, joint optimization with a physics engine, and sparse-reward reinforcement learning, enabling cross-task motion prior transfer and unconditional high-fidelity motion generation. Experiments demonstrate that the proposed representation significantly improves motion naturalness and temporal smoothness on challenging tasks—including head-mounted device tracking and non-uniform temporal interpolation—outperforming existing latent representation approaches in both qualitative and quantitative evaluations.
This work addresses the challenge of capturing the inherent uncertainty in human motion within football scenarios through 3D skeletal representations. To this end, the authors propose a self-supervised representation learning framework that leverages future motion prediction as a proxy task. The approach innovatively introduces a conditional module in 3D Euclidean space to model the multimodal probability distribution of discretized future motions, thereby explicitly capturing multiple plausible motion trajectories. Experimental results demonstrate that the learned representations significantly improve prediction accuracy on large-scale football tracking data and exhibit strong cross-task generalization capabilities across diverse downstream tasks.
This work addresses two key challenges in 3D human scanning and motion modeling: (1) fast, robust embedding estimation for unregistered meshes, and (2) geometrically faithful modeling of pose parameter spaces. We propose two self-supervised frameworks—VariShaPE and MoGeN. VariShaPE introduces a novel *variational manifold* architecture for shape parameter estimation, enabling millisecond-level encoding of unregistered meshes without correspondence. MoGeN embeds SMPL’s latent pose space into a higher-dimensional Euclidean space, yielding the first linearly interpolatable motion representation—supporting zero-computation-cost interpolation, extrapolation, and cross-action transfer. Both methods are fully data-driven and require no manual annotations. Quantitative and qualitative evaluations demonstrate significant improvements in generation fidelity and editing flexibility. The resulting latent spaces are computationally efficient and geometrically consistent, establishing a robust foundation for real-time human modeling and animation.
This work addresses the limitation of existing vision-language-action (VLA) models in lacking explicit motion priors, which hinders their ability to jointly model action temporal dynamics and multimodal alignment across diverse robotic embodiments. To overcome this, the authors propose a two-stage training framework: first, an unconditional action trajectory pretraining stage learns a decoupled action module that captures universal temporal motion structures transferable across robot morphologies; second, this motion prior is integrated into VLA joint training. The approach incorporates a history compression mechanism for efficient temporal context generation and employs a lightweight flow-matching-based encoder-decoder architecture with decoder reuse and early latent distillation. Evaluated on 13 cross-embodiment tasks, the method significantly improves convergence speed and success rates—particularly in data-scarce real-world scenarios—and demonstrates that scaling action data effectively enhances downstream generalization.
This work addresses the limited generalization of existing models in multimodal action prediction due to inadequate temporal representation, particularly when encountering unseen action sequences. Inspired by the mirror neuron system, the authors propose DMBN-PTE, a novel framework that integrates Conditional Neural Processes (CNP), a Deep Multimodal Fusion Network (DMBN), and Positional Temporal Encoding (PTE). The model leverages self-supervised learning to reconstruct visuomotor signals from partial observations, enabling robust long-horizon action prediction. A key innovation lies in the introduction of positional temporal encoding, which substantially enhances the model’s capacity to capture temporal dynamics. Experimental results demonstrate that DMBN-PTE achieves superior generalization on unseen action sequences, offering an effective solution for long-term action prediction in robotic applications.
This work addresses the challenge of learning generalizable action representations from visual dynamics to enhance the sample efficiency and generalization of world models in low-data regimes. To this end, the authors propose SCAR, a framework that leverages a pretrained generative backbone and jointly trains inverse and forward dynamics models to encode actions as disentangled latent factors governing controllable visual changes. Disentanglement and cross-embodiment transferability are achieved through Gaussian prior regularization and adversarial invariance constraints. Empirical evaluation demonstrates that SCAR significantly improves both sample efficiency and transfer performance of world models on the Procgen and Robotwin benchmarks, enabling effective generalization across tasks and embodiments.