synthesize video from text

Design, build, and evaluate models and pipelines that synthesize moving-image sequences from natural-language descriptions, producing temporally coherent, variable-length videos that reflect the input text. Work encompasses conditioning on textual or class labels, sampling diverse motion and viewpoints, enforcing temporal coherence and duration, applying post-processing and filtering, and measuring realism and fidelity to the prompt.

synthesizevideofromtext

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.

caption qualitydata distributiontext-to-video generation

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification

Nov 22, 2024
SP
S. P. Sharan
🏛️ The University of Texas at Austin

Current text-to-video generation models (e.g., Sora, Gen-3) are predominantly evaluated using metrics emphasizing visual quality and motion smoothness, while neglecting temporal fidelity and text-video alignment—critical requirements for safety-critical applications. To address this gap, we propose NeuS-V, the first quantitative evaluation framework grounded in neural-symbolic formal verification. Our method comprises three key components: (1) automatic compilation of natural language prompts into temporal logic (TL) specifications; (2) symbolic modeling of videos as finite-state automata; and (3) rigorous formal verification via model checking. To support evaluation of temporal complexity, we construct the first synthetic prompt dataset explicitly designed for varying temporal intricacy. Experiments demonstrate that NeuS-V achieves over fivefold higher correlation with human judgment compared to existing metrics and, for the first time, systematically exposes severe temporal reasoning failures of state-of-the-art models under temporally complex prompts.

Addressing gaps in current video generation evaluation metricsAssessing text-to-video alignment in synthetic video modelsEvaluating temporal fidelity using neuro-symbolic formal verification

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

Aug 15, 2024
JF
Jiasong Feng
🏛️ AI Research | Sun Yat-sen University

Existing text-to-video (T2V) models struggle to generate long-duration, motion-rich videos with strong temporal coherence, particularly in modeling implicit temporal logic within prompts and enabling frame-level fine-grained text guidance. To address this, we propose the Cross-Frame Text-Guided Module (CTGM), the first framework integrating a Temporal Information Injector (TII), a Temporal Affinity Refiner (TAR), and a Temporal Feature Booster (TFB) to achieve frame-specific text conditioning and dynamic temporal alignment. Built upon a diffusion-based architecture, our method incorporates latent-space temporal injection, cross-frame text-feature correlation recalibration, and implicit temporal consistency enhancement. Extensive experiments across multiple benchmarks demonstrate significant improvements in motion coherence and semantic fidelity. Both qualitative and quantitative evaluations consistently surpass state-of-the-art methods. The code, pre-trained models, and demonstration videos are publicly released.

Enhancing dynamic and consistent video generation from textImproving temporal logic comprehension in video synthesisSupporting both text-to-video and image-to-video generation tasks

Motion Prompting: Controlling Video Generation with Motion Trajectories

Dec 03, 2024
DG
Daniel Geng
🏛️ University of Michigan | Google DeepMind | Brown University

Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.

Control video generation using motion trajectories instead of text promptsEncode flexible motion representations for object-specific or global scene motionTranslate high-level user requests into detailed motion prompts for diverse applications

Latest Papers

What's happening recently
View more

Existing video-language models struggle to accurately describe fine-grained details—such as subjects, scenes, motion dynamics, and cinematographic parameters—in professional videos like films and advertisements. This work proposes the CHAI framework, which integrates structured visual primitives defined by expert creators with a critical human-AI collaboration mechanism: the model generates captions, which domain experts then critique and refine, thereby efficiently producing high-quality annotations and strong supervisory signals. By leveraging structured description schemas, supervised fine-tuning (SFT), direct preference optimization (DPO), and inference-time expansion, the approach substantially enhances Qwen3-VL’s performance with minimal expert intervention, outperforming Gemini-3.1-Pro. Furthermore, it successfully re-annotates a large-scale professional video dataset, significantly improving the fine-grained controllability of video generation models such as Wan on complex prompts up to 400 words in length.

human-AI oversightprecise captioningprofessional video understanding

Factorized Video Generation: Decoupling Scene Construction and Temporal Synthesis in Text-to-Video Diffusion Models

Dec 18, 2025
MH
Mariam Hassan
🏛️ École Polytechnique Fédérale de Lausanne

Existing text-to-video (T2V) diffusion models generate visually high-fidelity videos but often fail in scene construction and exhibit motion artifacts due to semantically ambiguous or logically inconsistent initial frames. To address this, we propose a semantic anchor-frame-driven factorized video generation framework that decouples T2V into three sequential stages: (1) LLM-guided semantic reasoning for coherent scene specification; (2) high-fidelity anchor frame generation via a text-to-image (T2I) model; and (3) temporal evolution modeling exclusively by a video diffusion model. This paradigm achieves the first explicit separation of scene construction and motion modeling responsibilities, significantly improving logical consistency and motion controllability. Integrated with prompt rewriting, visual anchoring sampling, and lightweight fine-tuning, our method achieves state-of-the-art performance on both T2V CompBench and VBench2, while reducing sampling steps by 70% without compromising visual quality.

Decoupling scene construction from temporal synthesis in video generationImproving logical coherence and compositional accuracy in text-to-video modelsResolving semantic inconsistencies in initial frames of generated videos

Existing text-to-video models struggle to generate physically consistent dynamic content due to their reliance on implicit temporal modeling. This work proposes a dual-engine agent framework that, for the first time, leverages executable Blender code as a procedural intermediate representation. In this approach, an encoding agent generates programs describing scene composition and temporal evolution; a simulation engine executes these programs to produce deterministic spatiotemporal drafts, which are then refined by a video generation engine into photorealistic outputs. By decoupling procedural reasoning from high-fidelity rendering, the method significantly enhances controllability, interpretability, and physical consistency. Trained on a newly curated VideoCoCo-3K dataset comprising draft-instruction-target triplets, the model achieves state-of-the-art performance with scores of 0.558 on PhyGenBench and 77.88 on VBench-2.0.

executable representationphysically-consistent video generationspatiotemporal consistency

Real-world multimodal video data suffers from high acquisition costs and limited diversity, hindering the training of large-scale multitask video understanding models. To address this challenge, this work proposes the first unified synthetic data generation framework capable of automatically producing unlimited, multitask-compatible multimodal video data. The approach introduces a visual question answering (VQA)-based fine-tuning strategy that replaces conventional caption- or instruction-based supervision with structured question-answer pairs to enhance the model’s visual reasoning and localization capabilities. Remarkably, models trained exclusively on this synthetic data achieve performance on par with or even surpassing fully supervised baselines on three distinct tasks—video object counting, video question answering, and video segmentation—demonstrating strong generalization and effectiveness across real-world benchmarks.

data annotationlarge language modelsmultimodal video understanding

Existing generative video models struggle to achieve human-centric camera control aligned with cinematic language, often producing random camera trajectories, spatial inconsistencies, and insufficient focus on the human subject. This work proposes a human-centric camera parameterization method that formalizes cinematic composition principles into computable, human-relative camera parameters for the first time. It introduces a domain-specific language (DSL) that coordinates with a multimodal large language model to map natural language instructions and human motion into cinematic keyframe shots, followed by deterministic interpolation to generate smooth, continuous camera trajectories. Evaluated on a newly curated dataset of 34K text–motion–camera aligned samples, the approach significantly outperforms existing methods on composition-oriented metrics, enabling controllable and aesthetically cinematic human-centric video generation.

camera controlcinematographic framinggenerative video models

Hot Scholars

YH

Yicong Hong

Adobe Research
Video GenerationWorld ModelsEmbodied AI
SJ

Simon Jenni

Adobe Research
computer visionmachine learningdeep learningunsupervised learning
AD

Angela Dai

Technical University of Munich
Computer GraphicsComputer Vision
WW

Weiping Wang

School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security