flow-based motion synthesis

Designs, implements, and evaluates models and pipelines that generate, reconstruct, or manipulate time-varying motion trajectories and kinematic sequences from conditioning inputs (such as video or other signals) using flow-based, autoregressive, diffusion, and physics-driven techniques. This work also builds methods for physics-aware trajectory synthesis and structure-from-motion, and defines quantitative motion-quality metrics and evaluation procedures to measure realism, fidelity, and adherence to desired dynamics.

flow-basedmotionsynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.21
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$199K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing video diffusion models struggle to generate realistic and controllable videos due to the scarcity of training samples exhibiting highly dynamic or precisely controlled motion. To address this limitation, this work proposes DynaVid, a two-stage diffusion framework that decouples motion from appearance. The first stage leverages synthetic optical flow—generated via computer graphics—as a supervisory signal to learn controllable dynamics; the second stage conditions on this learned motion representation to synthesize high-fidelity video frames. By explicitly separating motion modeling from appearance generation, DynaVid significantly enhances both visual realism and motion controllability, particularly in challenging scenarios involving vigorous human actions or extreme camera movements, outperforming current state-of-the-art methods.

dynamic video generationmotion controllabilityoptical flow

This work addresses the longstanding “last mile” problem in video-based motion capture, where reconstructed motions often exhibit physical implausibility and artifacts that necessitate extensive manual correction for industrial applications such as film and gaming. The authors propose a production-oriented physics-aware motion refinement framework that enhances both single- and multi-person motion sequences through physics-based optimization while seamlessly integrating keyframe editing to allow animators to inject stylistic adjustments. Developed in close collaboration with professional animators, the method balances automation efficiency with artist control, significantly improving the physical plausibility and visual quality of motion data. This approach markedly reduces post-processing effort and is designed to fit directly into real-world animation pipelines.

motion artifactsmotion capturephysical realism

Motion Prompting: Controlling Video Generation with Motion Trajectories

Dec 03, 2024
DG
Daniel Geng
🏛️ University of Michigan | Google DeepMind | Brown University

Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.

Control video generation using motion trajectories instead of text promptsEncode flexible motion representations for object-specific or global scene motionTranslate high-level user requests into detailed motion prompts for diverse applications

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

Sep 24, 2025
CW
Chen Wang
🏛️ University of Pennsylvania | MIT | HKUST

Existing video generation models produce high-fidelity videos but often lack physical plausibility and 3D controllability. To address this, we propose a physics-anchored image-to-video generation framework. Our method introduces a generative physics network that explicitly models multi-material dynamics—including elastic bodies, granular media (e.g., sand), viscoelastic putty, and rigid bodies—alongside a spatiotemporal attention module to capture inter-particle interactions. We jointly optimize trajectory plausibility and visual quality via a composite loss incorporating physics-based constraints. Furthermore, we employ a diffusion model to synthesize physically consistent 3D point trajectories, which drive controllable video synthesis. Trained on 550K synthetic samples, our approach surpasses state-of-the-art methods in both physical plausibility and visual fidelity. It enables fine-grained dynamic editing guided by physical parameters (e.g., elasticity, friction) and external forces (e.g., gravity, impact), offering unprecedented control over physically grounded video generation.

Addressing limited 3D controllability in existing video generation methodsGenerating physics-grounded motion with parameter and force controlOvercoming lack of physical plausibility in video generation models

Latest Papers

What's happening recently
View more

MotionV2V: Editing Motion in a Video

Nov 25, 2025
RB
Ryan Burgert
🏛️ Google | Stony Brook University

Despite significant advances in video generation models regarding fidelity and temporal coherence, precise and controllable motion editing of existing videos remains challenging. This paper introduces a novel “motion editing” paradigm: first, extracting sparse motion trajectories from input videos to enable fine-grained, arbitrary-timestep trajectory modifications; second, constructing a motion counterfactual video dataset and designing a motion-conditioned video diffusion architecture to naturally propagate edited trajectories and re-render the video. Our approach unifies sparse trajectory editing with generative resynthesis for the first time, enabling high-fidelity, temporally consistent motion redirection. A user study (four-alternative forced choice) demonstrates that over 65% of participants prefer our results—significantly outperforming state-of-the-art methods.

Editing motion trajectories in existing videosEnabling timestamp-specific motion edits with natural propagationGenerating motion counterfactuals with identical content

Current image-to-video generation models struggle to accurately simulate mechanical motion governed by kinematic and geometric constraints, often exhibiting inconsistencies in rigidity preservation, component contact, and motion transmission. This work proposes MechVerse—the first benchmark dataset specifically designed for mechanical assembly scenarios—which systematically defines and quantifies mechanical motion consistency in video generation. The benchmark encompasses three levels of mechanism complexity and establishes a multi-tiered evaluation framework integrating synthetic data, structured prompts, standard video metrics, instruction-following scores, and human assessments of motion correctness. Experiments reveal that while state-of-the-art models maintain visual fidelity and temporal smoothness, they perform poorly in terms of mechanical plausibility, with error rates rising significantly as coupling complexity increases.

kinematic constraintsmechanical assembliesmotion correctness

Existing video editing methods struggle to globally translate the 3D trajectories of objects while preserving their relative motion structure and scene plausibility. To address this challenge, this work proposes TrajectoryMover, a generative model–based approach for controllable video editing. By constructing TrajectoryAtlas—a large-scale synthetic paired dataset—and incorporating 3D motion consistency constraints during fine-tuning, TrajectoryMover achieves, for the first time, holistic 3D trajectory transfer of objects in videos. The method significantly outperforms existing techniques while maintaining object identity, motion semantics, and visual realism, thereby establishing a new paradigm for high-fidelity, controllable video editing.

3D motiongenerative movementobject trajectory

This work proposes the first end-to-end differentiable framework capable of generating physically plausible dynamics in 3D scenes directly from text prompts, circumventing the need for expert knowledge or laborious parameter tuning characteristic of traditional physics simulation. By leveraging a video diffusion model to extract motion priors and introducing a learnable motion distillation loss, the method effectively decouples motion from appearance and geometric discrepancies. Notably, it achieves purely text-driven generation of physically consistent dynamics without requiring ground-truth trajectories or annotated videos. Evaluated across more than 30 diverse scenarios—including elastic solids, metals, foams, granular materials, and both Newtonian and non-Newtonian fluids—the approach significantly outperforms existing methods in producing realistic and physically coherent motion.

3D object dynamicsmotion simulationnatural language prompting

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
YL

Yebin Liu

Professor, Tsinghua University
Computer GraphicsComputational Photography3D VisionDigital Humans
KF

Ke Fan

Fudan University
Machine LearningDeep Learning
HW

He Wang

Assistant Professor of Computer Science, Peking University
Embodied AIComputer VisionRobotics