Score
Designs and builds systems and datasets that generate synthetic video sequences (including paired, stylized, or otherwise transformed variants) together with corresponding labels or correspondences for training and evaluation. This work includes pipelines for reverse data synthesis and large-scale automated generation of video training pairs with controlled appearance and temporal variations to replace or augment real video data.
Existing text-to-video generation models struggle to support complex, fine-grained user control. This paper presents a systematic survey of controllable video generation, proposing a three-tier taxonomy—uniconditional, multiconditional, and general controllable—categorized by control signal modality. It unifies the theoretical foundations and fusion mechanisms for multimodal conditioning (e.g., camera motion, depth maps, human pose) within diffusion-based video generation. Methodologically, we design a conditional guidance architecture tailored for video diffusion models, enabling joint spatiotemporal integration of textual and non-textual control signals during denoising. Key contributions include: (1) the first comprehensive taxonomy for controllable video generation; (2) an open-source, unified evaluation benchmark and codebase; and (3) significantly enhanced controllability and practical applicability of AIGC in real-world scenarios.
This work addresses scene-aware multi-step visual instruction generation: given an initial scene image and a textual step-by-step instruction, the task is to generate a semantically coherent and environment-adapted sequence of instructional images. Methodologically, we propose a video-driven framework for automatic large-scale visual instruction data construction; introduce ShowHowTo—the first scene-conditioned video diffusion model—integrating instructional video mining, scene-conditioned generation, and multi-step alignment learning; and establish a comprehensive evaluation protocol spanning step-, scene-, and task-level dimensions. Trained on 0.6M high-quality image-text sequences, our approach achieves state-of-the-art performance across three accuracy metrics. All code, datasets, and models are publicly released.
This work addresses the challenge of high-quality multimodal video generation and editing by proposing a unified multimodal foundation model architecture. Methodologically, it introduces variable-aspect-ratio 1080p video latent-space modeling, cross-modal alignment training across text, image, video, and audio modalities, efficient tokenization, large-scale parallel training and inference optimization, and a rigorously quality-controlled data curation strategy coupled with a novel evaluation protocol. Key contributions include the first 30-billion-parameter video generation model supporting long-horizon generation (73K tokens, i.e., 16 seconds at 16 fps), instruction-driven precise editing, user-provided image personalization, and synchronized audio-video synthesis. The model achieves state-of-the-art performance across five benchmarks: text-to-video, video personalization, video editing, video-to-audio, and text-to-audio—demonstrating substantial improvements in temporal coherence and semantic controllability.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
Controllable human video generation is hindered by the scarcity of real-world data, particularly for rare identities and complex motion scenarios. This work proposes a unified diffusion-based framework that systematically investigates, for the first time, the synergistic mechanisms between synthetic and real data in human-centric video generation. It reveals their complementary roles and introduces an efficient synthetic sample selection strategy to enhance training. The proposed approach significantly improves motion realism, temporal coherence, and identity fidelity in generated videos, establishing a new paradigm for building data-efficient and generalizable controllable video generation models.
This work addresses the appearance inconsistency in autoregressive video generation when the camera revisits previously observed locations, a problem caused by overwriting of contextual cache that breaks 3D scene consistency. The authors propose a training-free method that leverages pose and depth information from a 3D engine to retrieve historical latent patches via spatiotemporal correspondences and inject them into the key-value (KV) cache. Additionally, they introduce a geometry-aware attention bias based on depth reprojection to align features geometrically. This approach is the first to integrate spatiotemporal correspondences without additional training, effectively closing the visual loop. It significantly outperforms existing training-free baselines on revisit trajectories in the TartanAir and TartanGround datasets while preserving high-quality video generation.
This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.
Existing video stylization methods are often hindered by content leakage, limited training data, and poor adaptability to long videos, leading to style drift and motion distortion. To address these challenges, this work proposes a text-driven video-to-video generation framework that introduces a novel automated reverse synthesis pipeline to construct V-Style20k, a large-scale video stylization dataset. The approach further incorporates an initialization-following mechanism and a sliding-window inference strategy, enabling high-quality style transfer for videos of arbitrary length. Experimental results demonstrate that the proposed method achieves superior performance across diverse artistic styles, effectively mitigating style drift and motion artifacts while matching the quality of leading closed-source solutions.