video generation

Designs, builds, or evaluates models and end-to-end pipelines that synthesize moving-image sequences (frames with temporal continuity) from inputs such as text, images, audio, or latent codes, controlling appearance, motion, timing, and consistency across frames. Work includes model architecture and training, dataset curation, temporal and perceptual quality metrics, and generation post-processing such as rendering, upsampling, and compression.

videogeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.66
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$192K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Pathways on the Image Manifold: Image Editing via Video Generation

Nov 25, 2024
NR
Noam Rotstein
🏛️ Technion - Israel Institute of Technology

Existing image diffusion models exhibit low accuracy in complex text-guided editing and often degrade critical content of the original image. To address this, we reformulate static image editing as a temporal evolution process and, for the first time, leverage a pre-trained image-to-video diffusion model to synthesize a manifold-continuous transition path—from source to edited image—along an implicit temporal trajectory. This temporal consistency enforces spatial semantic coherence without requiring fine-tuning or additional training. Our approach integrates manifold-constrained optimization and implicit temporal modeling to jointly preserve editing fidelity and structural/identity integrity of the input. Evaluated on text-driven image editing, our method achieves state-of-the-art performance, outperforming leading approaches both quantitatively (e.g., higher CLIP-Score, lower LPIPS) and qualitatively (e.g., sharper details, better semantic alignment, and stronger identity preservation).

Improving accuracy in complex image editsPreserving key elements of original imagesUsing video models for consistent image editing

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

This work addresses the challenge of high-quality multimodal video generation and editing by proposing a unified multimodal foundation model architecture. Methodologically, it introduces variable-aspect-ratio 1080p video latent-space modeling, cross-modal alignment training across text, image, video, and audio modalities, efficient tokenization, large-scale parallel training and inference optimization, and a rigorously quality-controlled data curation strategy coupled with a novel evaluation protocol. Key contributions include the first 30-billion-parameter video generation model supporting long-horizon generation (73K tokens, i.e., 16 seconds at 16 fps), instruction-driven precise editing, user-provided image personalization, and synchronized audio-video synthesis. The model achieves state-of-the-art performance across five benchmarks: text-to-video, video personalization, video editing, video-to-audio, and text-to-audio—demonstrating substantial improvements in temporal coherence and semantic controllability.

Develop high-quality 1080p HD video generation.Enable precise instruction-based video editing.Generate personalized videos using user images.

Magic Fixup: Streamlining Photo Editing by Watching Dynamic Videos

Mar 19, 2024
HA
Hadi Alzayer
🏛️ Adobe | University of Maryland, College Park

This work addresses the challenge of synthesizing high-fidelity, photorealistic images from coarse layout edits. To mitigate second-order artifacts—including illumination mismatch, missing shadows, and physically implausible object interactions—the authors propose a diffusion-based inpainting method leveraging video temporal modeling. The method introduces a novel dual-motion modeling mechanism—optical-flow-guided warping coupled with hierarchical feature injection—supervised by paired video frames, enabling joint optimization of layout alignment, illumination consistency, and physically grounded object interactions. By integrating a pre-trained diffusion model, layout-constrained fine-tuning, and a dynamically constructed video dataset, the approach achieves fine-grained detail transfer and multi-factor coherent generation. Experiments demonstrate significant improvements in output photorealism, geometric consistency, and scene plausibility, while preserving object identity and texture fidelity.

Generates photorealistic images from coarse editsTransfers fine details while adapting to new lightingUses video supervision for realistic object interactions

Latest Papers

What's happening recently
View more

Existing text-to-video models struggle to generate physically consistent dynamic content due to their reliance on implicit temporal modeling. This work proposes a dual-engine agent framework that, for the first time, leverages executable Blender code as a procedural intermediate representation. In this approach, an encoding agent generates programs describing scene composition and temporal evolution; a simulation engine executes these programs to produce deterministic spatiotemporal drafts, which are then refined by a video generation engine into photorealistic outputs. By decoupling procedural reasoning from high-fidelity rendering, the method significantly enhances controllability, interpretability, and physical consistency. Trained on a newly curated VideoCoCo-3K dataset comprising draft-instruction-target triplets, the model achieves state-of-the-art performance with scores of 0.558 on PhyGenBench and 77.88 on VBench-2.0.

executable representationphysically-consistent video generationspatiotemporal consistency

Existing methods struggle to simultaneously achieve high appearance fidelity, natural motion dynamics, and controllable camera viewpoints in human video synthesis when multi-view data are limited. This work proposes an “image-first” generation paradigm: it first leverages a pre-trained image generation model to learn a high-quality human appearance prior, then integrates SMPL-X pose conditioning with a pre-trained video diffusion model. Through a training-free temporal optimization strategy, the approach enables pose- and viewpoint-controllable, high-fidelity video synthesis. By effectively decoupling appearance modeling from temporal consistency, the method significantly enhances both visual quality and controllability. The authors also release a standardized human dataset and the corresponding synthesis model to support future research.

appearance modelingcamera viewpointhuman video generation

This work addresses the lack of standardized, automated evaluation methods for design animation video generation, which hinders objective assessment of generation quality under structured constraints. To bridge this gap, we propose the first multidimensional automatic evaluation framework tailored specifically for design animations. Leveraging computer vision and video analysis techniques, the framework quantifies key generative attributes across four dimensions: layout fidelity, motion correctness, temporal consistency, and content fidelity. Operating without human intervention, it delivers an objective and reproducible benchmark that enables fair comparison among diverse generative models and supports sustained progress in the field.

compositional fidelitydesign video generationevaluation framework

Existing image generation models struggle to capture the temporal evolution of the visual world, often failing to maintain cross-frame identity, spatial relationships, and causal ordering. To address this limitation, this work introduces ImageTime, a benchmark that evaluates models’ ability to generate temporally coherent images under sequential instructions using a four-keyframe protocol—comprising initial, action-onset, transition, and final states—thereby abstracting away video-level dynamics to focus on logical consistency between static frames. We propose a novel diagnostic evaluation framework centered on spatiotemporal consistency as a probing mechanism, featuring a hierarchical task structure and structured state predicates. Leveraging a VLM-as-judge paradigm enables interpretable scoring and failure attribution. Through multi-stage state definitions, temporal constraint modeling, and causal violation detection—augmented by GPT-5.5–driven automated assessment—our approach systematically uncovers the capability boundaries, failure modes, and concept drift phenomena of state-of-the-art models in temporal visual consistency.

coherent time imaginationimage generationspatiotemporal modeling

This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.

caption qualitydata distributiontext-to-video generation

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
HQ

Haotong Qin

ETH Zürich
TinyMLModel CompressionComputer VisionDeep Learning
ZZ

Zhaoxiang Zhang

Institute of Automation, Chinese Academy of Sciences
Computer VisionPattern RecognitionBiologically-inspired Learning
WM

Weijia Mao

a phd student at National University of Singapore
computer vision3D generation and reconstruction
YW

Yaohui Wang

Research Scientist, Shanghai AI Laboratory | Inria
Machine LearningDeep Generative ModelsVideo Generation