Score
Designs and implements models and pipelines that translate, synthesize, or generate video sequences conditioned on proxy representations or dynamics—e.g., performing video-to-video translation, proxy-conditioned video generation, and motion- or dynamics-conditioned video synthesis. This includes constructing systems that produce synthetic proxy videos, encode local or per-pixel rigid-body motion, absorb appearance and occlusion effects, and train or fine-tune video generative models (such as diffusion models) on synthetic proxy data.
Existing text-to-video generation models struggle to support complex, fine-grained user control. This paper presents a systematic survey of controllable video generation, proposing a three-tier taxonomy—uniconditional, multiconditional, and general controllable—categorized by control signal modality. It unifies the theoretical foundations and fusion mechanisms for multimodal conditioning (e.g., camera motion, depth maps, human pose) within diffusion-based video generation. Methodologically, we design a conditional guidance architecture tailored for video diffusion models, enabling joint spatiotemporal integration of textual and non-textual control signals during denoising. Key contributions include: (1) the first comprehensive taxonomy for controllable video generation; (2) an open-source, unified evaluation benchmark and codebase; and (3) significantly enhanced controllability and practical applicability of AIGC in real-world scenarios.
Existing video generation models struggle to precisely control complex, physically plausible dynamics and object interactions through text alone. This work proposes a training-free proxy-conditioned video generation approach that leverages coarse-grained proxy videos—derived from physics simulations or real-world recordings—as dynamic priors to synthesize novel content with temporally coherent motion guided by text prompts. Built upon a pretrained video diffusion model, the method integrates proxy dynamics and textual semantics through latent-space inversion of the proxy video, region-aware latent noise injection, and a Stochastic Flow Relaxation (SFR) mechanism. Experiments demonstrate that the proposed approach significantly outperforms current video editing and motion transfer techniques in both dynamic fidelity and alignment with textual descriptions.
Current video generation models trained on a single subject image struggle to simultaneously achieve motion diversity and compositional generalization of subject concepts, while requiring frequent retraining with new video data—entailing high interactive overhead. To address this, we propose a fine-tuning-free, text-image jointly driven personalized video generation framework. Our method introduces a plug-and-play lightweight subject learning module, coupled with a dynamic weighted video sampling strategy that preserves motion priors in early denoising steps and emphasizes appearance reconstruction in later steps. We further incorporate parameter-efficient Video Diffusion Model (VDM) updates, phased weight scheduling, and cross-modal concept alignment. Extensive experiments demonstrate significant improvements over state-of-the-art methods across multiple benchmarks, achieving superior motion naturalness, subject fidelity, and flexibility in compositional concept generation. The source code is publicly available.
Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.
To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
This work proposes a novel approach to 6-DoF pose tracking of textureless, transparent, reflective, or deformable objects in monocular videos without relying on additional inputs such as 3D models, depth maps, or object masks. Framing pose tracking as a video-to-video translation task, the method requires only the raw input video and a single annotated pixel in the first frame. A fine-tuned video diffusion model generates a proxy video with known geometry and appearance, enabling robust pose estimation via classical algorithms. The framework makes no assumptions about object identity, boundaries, or global rigidity, and achieves state-of-the-art accuracy using only synthetic data for fine-tuning. It effectively handles complex materials, occlusions, and deformations, and generalizes successfully to diverse applications including face tracking, camera pose estimation, and challenging real-world scenes.
为了解决动作条件视频模型数据获取难题,本文提出基于虚幻引擎的两阶段合成数据生成管道,生成大规模动作条件多视角视频。
This study addresses the issues of appearance attribute leakage and mode collapse in video generation caused by regressing reference videos. To overcome these limitations, this work proposes a motion customization framework based on stochastic optimal control (SOC). By integrating SOC into diffusion models to guide dynamic generation and designing a timestep-adaptive motion cost function, the approach achieves precise motion control without requiring explicit rewards. The proposed method effectively mitigates content leakage while preserving high motion fidelity and diversity, ultimately improving training efficiency by 2.5 times.
SpatialCrafter通过引入全局3D代理和两阶段框架,解决了基于视频扩散模型的图像到场景生成中的随机幻觉、长期漂移及3D一致性不佳的问题。
This study addresses the limitation of existing motion generation models that overlook the animation principle of exaggeration, resulting in character motions lacking expressiveness. We propose integrating this principle into diffusion- and flow matching-based motion generation frameworks. During training, exaggeration priors are injected via supervised fine-tuning. For inference, we introduce a novel mathematical formulation of Dynamic Movement Primitives (DMPs) to enable training-free exaggeration guidance without retraining. This approach significantly enhances the exaggeration and artistic expressiveness of generated motions while strictly preserving physical plausibility and the original motion intent. Ultimately, this work establishes a new paradigm for generating controllable and dynamically compelling character animations.