3d motion extraction

Design and build pipelines that take video frames (often monocular) and produce temporally consistent 3D motion representations such as joint trajectories, articulated meshes, or facial motion sequences for single or multiple subjects. Implement and analyze algorithms that iteratively refine motion estimates—using coarse‑to‑fine and iterative generation strategies, motion encoders and feedback loops—and fuse complementary cues (e.g., mesh features and optical flow) via cross‑attention or other fusion mechanisms to improve accuracy and extract secondary signals like foot contacts or step counts.

3dmotionextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.18
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the common issue in monocular video-based human motion recovery—over-smoothing and kinematic inconsistency due to the lack of high-order temporal dynamics such as velocity and acceleration. To this end, the authors propose HTD-Refine, a post-processing framework that explicitly incorporates high-order temporal dynamics to enhance motion plausibility. At its core, PVA-Net leverages a temporal Transformer to jointly predict 2D joint positions, 3D velocities, and accelerations from monocular video, which are then integrated as soft constraints into a physics-inspired global trajectory optimization. The entire pipeline is end-to-end trainable and consistently outperforms existing methods across multiple in-the-wild benchmarks, effectively mitigating over-smoothing and jitter while recovering more accurate global trajectories and natural dynamic motion.

Dynamic ConsistencyHuman Motion RecoveryMonocular Videos

This work addresses the challenge of recovering high-quality, spatiotemporally consistent 4D dynamic objects from monocular videos, which is hindered by data scarcity and viewpoint ambiguity. The authors propose decomposing 4D synthesis into static 3D shape generation and motion reconstruction, introducing a canonical reference mesh to learn a compact implicit motion representation. A frame-wise Transformer predicts per-frame vertex trajectories, enabling geometrically consistent dynamic reconstruction. The method employs a feed-forward architecture with a scalable Transformer, supporting processing of sequences of arbitrary length. Evaluated on standard benchmarks as well as a newly constructed high-fidelity ground-truth dataset, the approach outperforms existing methods in both geometric fidelity and spatiotemporal consistency.

3D motion reconstruction4D synthesisdynamic objects

Existing 4D generation methods suffer from limitations in animation quality, temporal consistency, and controllability, primarily due to reliance on computationally expensive dense representations or insufficient fine-grained spatiotemporal control. This work proposes ACT, a framework that leverages lightweight skeletons as structured representations and employs 3D point trajectories extracted from monocular videos as explicit motion guidance to enable topology-agnostic, precise skeletal animation control. The core innovation lies in a routing trajectory injector that integrates prior-guided hard routing, a global routing mechanism, and local window cross-attention to achieve robust mapping from trajectories to joints, thereby enhancing whole-body motion awareness and micro-temporal alignment. Experiments demonstrate that ACT significantly outperforms existing approaches in animation fidelity, temporal consistency, and controllability, enabling high-quality, long-duration 4D animation synthesis.

3D trajectories4D generationmotion controllability

Generating Continual Human Motion in Diverse 3D Scenes

Apr 04, 2023
AM
Aymen Mir
🏛️ University of Tübingen | Max Planck Institute for Informatics | Meta AI Research | University of California, Berkeley

Generating long-horizon, multi-action coherent human motions in 3D scenes remains challenging due to drift accumulation, action discontinuity, and poor scene adaptability. Method: We propose an animator-guided, scene-agnostic iterative generation framework. It establishes a target-centric canonical coordinate system to decouple path planning from motion transition, and employs motion decomposition modeling with coordinate-system reparameterization—enabling zero-shot deployment on pure motion-capture data without scene-aware annotations or fine-tuning. Contribution/Results: To our knowledge, this is the first method to generate drift-free, chained multi-action sequences (e.g., “grasp → sit → lean”) in diverse real-world scanned environments—including HPS, Replica, Matterport, and ScanNet—using only sparse keypoint constraints and a seed motion. Unlike existing 3D navigation approaches, ours requires no scene rendering, geometric encoding, or environment-specific training, achieving superior generalization, motion plausibility, and scene compatibility.

3D Human Motion SimulationAdaptability to Different 3D EnvironmentsCoherence and Naturalness

Latest Papers

What's happening recently
View more

Traditional marker-based motion capture is costly and yields limited-scale, low-diversity datasets, hindering large-scale human motion modeling. This work proposes the first end-to-end scalable pipeline that automatically extracts 3D body and facial motions from in-the-wild internet videos and generates corresponding semantic textual descriptions, enabling the construction of high-quality, annotated motion datasets without controlled environments. By integrating monocular motion capture, video–language understanding, 3D pose and facial action estimation, and text generation, the method substantially enhances the flexibility and scalability of motion data acquisition. Motion reconstruction and generation models trained on this dataset achieve performance comparable to those trained on conventional motion capture data and demonstrate strong cross-dataset generalization capabilities.

3D motionhuman motion datasetin-the-wild

This work addresses the challenges of slow inference, low motion quality, and poor text-motion alignment in general category-agnostic 3D animation generation. The authors propose a feedforward generative framework that renders a rigged static 3D model into multi-layer images as input to a video generation model. By leveraging keypoint tracking to extract 2D joint motions and subsequently lifting them into 3D space to drive character animation, the method achieves efficient and natural motion synthesis. A key innovation lies in modeling 3D joint motion as transformations within a two-dimensional subspace, which significantly enhances both computational efficiency and motion naturalness. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods in terms of text-motion alignment, animation quality, and inference speed.

3D asset productioncategory-agnostic 3D animationinference speed

This work addresses the challenge in text-driven 3D human motion editing of simultaneously preserving the stylistic characteristics of the source motion while accurately adhering to textual instructions. To this end, the authors propose a dual-axis anchored Transformer architecture that separately extracts features along the joint and temporal dimensions and integrates them through a cross-axis fusion module for joint modeling. A novel auxiliary task—joint-level motion difference prediction—is introduced, leveraging Soft-DTW distance regression to guide the model in identifying critical joints and time steps requiring modification. Integrated within a diffusion-based generative framework and evaluated on the newly curated MotionFix dataset, the method significantly enhances both semantic alignment with input text and fidelity to the original motion structure, achieving state-of-the-art performance.

3D human motionjoint-wise editingmotion preservation

Existing methods for 3D multi-person motion prediction often generate skeletal sequences directly from noise, which frequently leads to structural inconsistencies and unreliable early-stage interactions. To address these limitations, this work proposes a prior-guided residual flow matching framework. The approach first leverages a deterministic coarse-grained motion prior to construct a residual conditional flow, thereby simplifying the generative objective. It further introduces a dynamic cross-interaction mechanism that enables temporally aligned multi-agent information synchronization during integration. Additionally, a decoupled joint-motion bidirectional fusion architecture is incorporated to preserve fine-grained motion consistency. Evaluated on multiple benchmark datasets, the proposed method achieves state-of-the-art performance, significantly improving prediction accuracy over existing approaches.

3D multi-person motion predictioncross-agent interactionsmotion residuals

This study addresses the challenge that acquiring 3D motion from monocular videos hinders the application of motion-language models. We propose a plug-and-play 2D motion interface that leverages cross-modal alignment and a real-video adapter to enable zero-shot adaptation of 3D-pretrained models to 2D inputs without architectural modifications. This approach effectively circumvents data bottlenecks by eliminating the need for model retraining. Evaluated on a newly constructed real-world video benchmark, our method achieves performance comparable to 3D-input baselines and significantly outperforms training from scratch. By overcoming limitations inherent to monocular video scenarios, this work establishes an efficient paradigm for the generalized deployment of motion understanding models, facilitating broader practical applications without reliance on scarce 3D motion data.

2D Motion InterfaceMonocular VideoMotion Language Models

Hot Scholars

ZL

Ziyuan Liu

Unknown affiliation
RoboticsManipulation and GraspingComputer VisionMachine Learning
JC

Jaehyun Choi

PhD Candidate @ KAIST
Dataset CondensationDomain Adaptation / Domain GeneralizationImage / Video Generation
RL

Ronghui Li

Tsinghua University
Human InteractionMotion GenerationDigital HumanComputer Vision
GH

Gyojin Han

KAIST
Deep LearningComputer Vision
CL

Changjian Li

Assistant Professor at University of Edinburgh
Computer Graphics3D VisionGeometry Analysis and ProcessingMedical Image Analysis