proxy video synthesis

Designs and implements models and pipelines that translate, synthesize, or generate video sequences conditioned on proxy representations or dynamics—e.g., performing video-to-video translation, proxy-conditioned video generation, and motion- or dynamics-conditioned video synthesis. This includes constructing systems that produce synthetic proxy videos, encode local or per-pixel rigid-body motion, absorb appearance and occlusion effects, and train or fine-tune video generative models (such as diffusion models) on synthetic proxy data.

proxyvideosynthesis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing video generation models struggle to precisely control complex, physically plausible dynamics and object interactions through text alone. This work proposes a training-free proxy-conditioned video generation approach that leverages coarse-grained proxy videos—derived from physics simulations or real-world recordings—as dynamic priors to synthesize novel content with temporally coherent motion guided by text prompts. Built upon a pretrained video diffusion model, the method integrates proxy dynamics and textual semantics through latent-space inversion of the proxy video, region-aware latent noise injection, and a Stochastic Flow Relaxation (SFR) mechanism. Experiments demonstrate that the proposed approach significantly outperforms current video editing and motion transfer techniques in both dynamic fidelity and alignment with textual descriptions.

controllable dynamicsdynamic fidelitymotion control

CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities

Aug 23, 2024
TW
Tao Wu
🏛️ Zhejiang University | Tencent | Polytechnic Institute

Current video generation models trained on a single subject image struggle to simultaneously achieve motion diversity and compositional generalization of subject concepts, while requiring frequent retraining with new video data—entailing high interactive overhead. To address this, we propose a fine-tuning-free, text-image jointly driven personalized video generation framework. Our method introduces a plug-and-play lightweight subject learning module, coupled with a dynamic weighted video sampling strategy that preserves motion priors in early denoising steps and emphasizes appearance reconstruction in later steps. We further incorporate parameter-efficient Video Diffusion Model (VDM) updates, phased weight scheduling, and cross-modal concept alignment. Extensive experiments demonstrate significant improvements over state-of-the-art methods across multiple benchmarks, achieving superior motion naturalness, subject fidelity, and flexibility in compositional concept generation. The source code is publicly available.

Action DiversityTraining EfficiencyVideo Generation

Motion Prompting: Controlling Video Generation with Motion Trajectories

Dec 03, 2024
DG
Daniel Geng
🏛️ University of Michigan | Google DeepMind | Brown University

Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.

Control video generation using motion trajectories instead of text promptsEncode flexible motion representations for object-specific or global scene motionTranslate high-level user requests into detailed motion prompts for diverse applications

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Aug 27, 2024
XW
Xiaojuan Wang
🏛️ University of Washington | Google | UC Berkeley

To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.

Adapting image-to-video modelsDual-directional diffusion sampling processGenerating video sequences between keyframes

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

Latest Papers

What's happening recently
View more

This work proposes a novel approach to 6-DoF pose tracking of textureless, transparent, reflective, or deformable objects in monocular videos without relying on additional inputs such as 3D models, depth maps, or object masks. Framing pose tracking as a video-to-video translation task, the method requires only the raw input video and a single annotated pixel in the first frame. A fine-tuned video diffusion model generates a proxy video with known geometry and appearance, enabling robust pose estimation via classical algorithms. The framework makes no assumptions about object identity, boundaries, or global rigidity, and achieves state-of-the-art accuracy using only synthetic data for fine-tuning. It effectively handles complex materials, occlusions, and deformations, and generalizes successfully to diverse applications including face tracking, camera pose estimation, and challenging real-world scenes.

6-DoF pose trackingdeformable objectsmonocular video

This study addresses the issues of appearance attribute leakage and mode collapse in video generation caused by regressing reference videos. To overcome these limitations, this work proposes a motion customization framework based on stochastic optimal control (SOC). By integrating SOC into diffusion models to guide dynamic generation and designing a timestep-adaptive motion cost function, the approach achieves precise motion control without requiring explicit rewards. The proposed method effectively mitigates content leakage while preserving high motion fidelity and diversity, ultimately improving training efficiency by 2.5 times.

content leakagegenerative collapsemotion customization

This study addresses the limitation of existing motion generation models that overlook the animation principle of exaggeration, resulting in character motions lacking expressiveness. We propose integrating this principle into diffusion- and flow matching-based motion generation frameworks. During training, exaggeration priors are injected via supervised fine-tuning. For inference, we introduce a novel mathematical formulation of Dynamic Movement Primitives (DMPs) to enable training-free exaggeration guidance without retraining. This approach significantly enhances the exaggeration and artistic expressiveness of generated motions while strictly preserving physical plausibility and the original motion intent. Ultimately, this work establishes a new paradigm for generating controllable and dynamically compelling character animations.

animation principlescharacter animationexaggeration

Hot Scholars

SB

Sagie Benaim

Assistant Professor, Hebrew University of Jerusalem
Computer VisionMachine Learning
XC

Xu Chen

Google
computer visionmachine learning
GE

George Eskandar

University of Stuttgart
Computer VisionDomain AdaptationGenerative AIAutonomous Driving
VC

Valerio Cambareri

Principal Software Engineer, Sony Depthsensing Solutions NV
Depth SensingTime-of-Flight ImagingComputational ImagingDeep Learning
DS

Dieter Schmalstieg

Alexander von Humboldt Professor of Visual Computing, University of Stuttgart
Augmented RealityVirtual RealityComputer GraphicsVisualization