learn multimodal action prior

Designs and trains probabilistic priors over motion and action sequences—multimodal, history‑conditioned, latent‑variable, and population-level—so they can be sampled, conditioned, or decoded into continuous kinematic trajectories or implicit neural representations for bodies, hands, or other articulated systems. This includes specifying loss formulations and estimators (e.g., IMLE/implicit MLE), enforcing kinematic and contact consistency, conditioning on recent history or semantic cues, and mapping latent codes to time-indexed pose or motion outputs.

learnmultimodalactionprior

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural Human Pose Prior

Jul 16, 2025
MH
Michal Heker
🏛️ Yoom

This work addresses the challenge of modeling probabilistic distributions of human poses represented as 6D rotations—i.e., elements on the SO(3) manifold. We propose a normalizing flow-based neural prior that explicitly respects the geometric structure of SO(3). Methodologically, we introduce the first integration of the RealNVP architecture with invertible Gram–Schmidt orthogonalization, enabling manifold-aware density estimation for 6D rotation representations. Our novel inverse Gram–Schmidt procedure ensures strict satisfaction of rotational constraints while preserving bijectivity and tractable likelihood evaluation. The resulting model exhibits strong representational capacity, training stability, and framework-agnostic compatibility. Experiments demonstrate substantial improvements in pose prior accuracy and cross-scenario generalization on motion capture and reconstruction tasks. Our approach establishes a new paradigm for probabilistic human motion modeling grounded in differential geometry and deep generative learning.

Integrating pose priors into motion capture pipelinesLearning flexible density over 6D rotation posesModeling neural prior over human body poses

This study investigates how humans efficiently recognize everyday activities from complex visual inputs and disentangles the critical motion information required for classification versus action reconstruction. By comparing three representation methods—Temporal Movement Primitives (TMP), Legendre polynomial coefficients, and autoencoders—the contributions of static body poses and temporal dynamics are analyzed across videos of 16 daily activities. The findings reveal that static poses alone suffice for high-accuracy activity classification, with nine key joints identified as most discriminative, whereas naturalistic action reconstruction fundamentally relies on temporal dynamics. TMP and Legendre coefficients achieve comparable, top performance in classification; however, only TMP generates perceptually natural motions, highlighting an essential decoupling between the informational demands of action recognition and action generation.

action recognitionmovement classificationmovement reconstruction

Versatile Physics-based Character Control with Hybrid Latent Representation

Mar 17, 2025
JB
Jinseok Bae
🏛️ Seoul National University

To address poor generalization, motion jitter, and difficulty in sharing priors in physics-based character motion control, this paper proposes a continuous-discrete hybrid latent representation architecture: continuous residual modeling captures fine-grained motion details, while Residual Vector Quantization (RVQ) enhances the expressiveness and diversity of discrete latent variables, effectively mitigating posterior collapse. The method integrates VAE extensions, joint optimization with a physics engine, and sparse-reward reinforcement learning, enabling cross-task motion prior transfer and unconditional high-fidelity motion generation. Experiments demonstrate that the proposed representation significantly improves motion naturalness and temporal smoothness on challenging tasks—including head-mounted device tracking and non-uniform temporal interpolation—outperforming existing latent representation approaches in both qualitative and quantitative evaluations.

Develops hybrid latent representation for physics-based character control.Enables diverse, smooth motions across multiple challenging tasks.Improves motion quality and expressiveness for sparse goal conditions.

This work addresses the challenge of capturing the inherent uncertainty in human motion within football scenarios through 3D skeletal representations. To this end, the authors propose a self-supervised representation learning framework that leverages future motion prediction as a proxy task. The approach innovatively introduces a conditional module in 3D Euclidean space to model the multimodal probability distribution of discretized future motions, thereby explicitly capturing multiple plausible motion trajectories. Experimental results demonstrate that the learned representations significantly improve prediction accuracy on large-scale football tracking data and exhibit strong cross-task generalization capabilities across diverse downstream tasks.

3D skeletonfuture motion predictionhuman motion uncertainty

Self Supervised Networks for Learning Latent Space Representations of Human Body Scans and Motions

Nov 05, 2024
EH
Emmanuel Hartman
🏛️ Florida State University | University of Houston

This work addresses two key challenges in 3D human scanning and motion modeling: (1) fast, robust embedding estimation for unregistered meshes, and (2) geometrically faithful modeling of pose parameter spaces. We propose two self-supervised frameworks—VariShaPE and MoGeN. VariShaPE introduces a novel *variational manifold* architecture for shape parameter estimation, enabling millisecond-level encoding of unregistered meshes without correspondence. MoGeN embeds SMPL’s latent pose space into a higher-dimensional Euclidean space, yielding the first linearly interpolatable motion representation—supporting zero-computation-cost interpolation, extrapolation, and cross-action transfer. Both methods are fully data-driven and require no manual annotations. Quantitative and qualitative evaluations demonstrate significant improvements in generation fidelity and editing flexibility. The resulting latent spaces are computationally efficient and geometrically consistent, establishing a robust foundation for real-time human modeling and animation.

Estimating latent space representations of body shapes and posesLearning geometry on latent space for motion approximationPerforming motion operations with limited computational cost

Latest Papers

What's happening recently
View more

This work addresses the limitation of existing vision-language-action (VLA) models in lacking explicit motion priors, which hinders their ability to jointly model action temporal dynamics and multimodal alignment across diverse robotic embodiments. To overcome this, the authors propose a two-stage training framework: first, an unconditional action trajectory pretraining stage learns a decoupled action module that captures universal temporal motion structures transferable across robot morphologies; second, this motion prior is integrated into VLA joint training. The approach incorporates a history compression mechanism for efficient temporal context generation and employs a lightweight flow-matching-based encoder-decoder architecture with decoder reuse and early latent distillation. Evaluated on 13 cross-embodiment tasks, the method significantly improves convergence speed and success rates—particularly in data-scarce real-world scenarios—and demonstrates that scaling action data effectively enhances downstream generalization.

action priorscross-embodimentmotion dynamics

This work addresses the limited generalization of existing models in multimodal action prediction due to inadequate temporal representation, particularly when encountering unseen action sequences. Inspired by the mirror neuron system, the authors propose DMBN-PTE, a novel framework that integrates Conditional Neural Processes (CNP), a Deep Multimodal Fusion Network (DMBN), and Positional Temporal Encoding (PTE). The model leverages self-supervised learning to reconstruct visuomotor signals from partial observations, enabling robust long-horizon action prediction. A key innovation lies in the introduction of positional temporal encoding, which substantially enhances the model’s capacity to capture temporal dynamics. Experimental results demonstrate that DMBN-PTE achieves superior generalization on unseen action sequences, offering an effective solution for long-term action prediction in robotic applications.

generalizationmultimodal action predictionneural processes

This work addresses the challenge of learning generalizable action representations from visual dynamics to enhance the sample efficiency and generalization of world models in low-data regimes. To this end, the authors propose SCAR, a framework that leverages a pretrained generative backbone and jointly trains inverse and forward dynamics models to encode actions as disentangled latent factors governing controllable visual changes. Disentanglement and cross-embodiment transferability are achieved through Gaussian prior regularization and adversarial invariance constraints. Empirical evaluation demonstrates that SCAR significantly improves both sample efficiency and transfer performance of world models on the Procgen and Robotwin benchmarks, enabling effective generalization across tasks and embodiments.

action representationcross-embodiment transferembodied intelligence

Hot Scholars

MH

Min-Hung Chen

Senior Research Scientist @ NVIDIA
Multimodal LearningVideo UnderstandingTransfer LearningComputer Vision
FE

Fu-En Yang

Research Scientist, NVIDIA Research
Artificial IntelligenceVision and LanguageMultimodal Learning