Score
Design, build, and evaluate models and pipelines that synthesize moving-image sequences from natural-language descriptions, producing temporally coherent, variable-length videos that reflect the input text. Work encompasses conditioning on textual or class labels, sampling diverse motion and viewpoints, enforcing temporal coherence and duration, applying post-processing and filtering, and measuring realism and fidelity to the prompt.
Text-to-video (T2V) generation remains hindered by weak text–visual alignment, poor long-term temporal coherence, and excessive computational overhead. This survey systematically traces the evolution of T2V models—from early GAN- and VAE-based approaches to contemporary diffusion-Transformer hybrids—analyzing architectural principles, pivotal advances, and drivers of paradigm shifts. It introduces the first unified framework integrating architecture design, training strategies, evaluation methodologies, and benchmark datasets; proposes a novel “perceptual alignment”-oriented evaluation paradigm; and conducts comprehensive, reproducible benchmarking of state-of-the-art models across standard datasets, exposing critical limitations of existing metrics while providing hardware specifications and hyperparameter configurations essential for replication. The study further identifies two key future directions: efficient lightweight generation and long-range spatiotemporal consistency.
This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
Current text-to-video generation models (e.g., Sora, Gen-3) are predominantly evaluated using metrics emphasizing visual quality and motion smoothness, while neglecting temporal fidelity and text-video alignment—critical requirements for safety-critical applications. To address this gap, we propose NeuS-V, the first quantitative evaluation framework grounded in neural-symbolic formal verification. Our method comprises three key components: (1) automatic compilation of natural language prompts into temporal logic (TL) specifications; (2) symbolic modeling of videos as finite-state automata; and (3) rigorous formal verification via model checking. To support evaluation of temporal complexity, we construct the first synthetic prompt dataset explicitly designed for varying temporal intricacy. Experiments demonstrate that NeuS-V achieves over fivefold higher correlation with human judgment compared to existing metrics and, for the first time, systematically exposes severe temporal reasoning failures of state-of-the-art models under temporally complex prompts.
Existing text-to-video (T2V) models struggle to generate long-duration, motion-rich videos with strong temporal coherence, particularly in modeling implicit temporal logic within prompts and enabling frame-level fine-grained text guidance. To address this, we propose the Cross-Frame Text-Guided Module (CTGM), the first framework integrating a Temporal Information Injector (TII), a Temporal Affinity Refiner (TAR), and a Temporal Feature Booster (TFB) to achieve frame-specific text conditioning and dynamic temporal alignment. Built upon a diffusion-based architecture, our method incorporates latent-space temporal injection, cross-frame text-feature correlation recalibration, and implicit temporal consistency enhancement. Extensive experiments across multiple benchmarks demonstrate significant improvements in motion coherence and semantic fidelity. Both qualitative and quantitative evaluations consistently surpass state-of-the-art methods. The code, pre-trained models, and demonstration videos are publicly released.
Existing video generation models rely heavily on text prompts, which lack precise spatiotemporal control over dynamic motion and complex action composition. To address this, we propose Motion Prompting—a novel conditioning framework that leverages variable-granularity motion trajectories (sparse/dense, object-level/global/temporal) to enable fine-grained control over camera/object motion, image interaction, motion transfer, and editing. Methodologically, we introduce the first trajectory encoder coupled with a spatiotemporal attention fusion architecture, complemented by motion-guided latent-space optimization and a semantic-driven motion prompt expansion mechanism that automatically maps high-level semantics into detailed motion signals. Quantitative evaluations and human studies across multiple tasks demonstrate significant improvements over state-of-the-art baselines. Generated videos exhibit enhanced physical plausibility and emergent behaviors, establishing a new paradigm for interactive video generation in embodied world modeling.
Existing video-language models struggle to accurately describe fine-grained details—such as subjects, scenes, motion dynamics, and cinematographic parameters—in professional videos like films and advertisements. This work proposes the CHAI framework, which integrates structured visual primitives defined by expert creators with a critical human-AI collaboration mechanism: the model generates captions, which domain experts then critique and refine, thereby efficiently producing high-quality annotations and strong supervisory signals. By leveraging structured description schemas, supervised fine-tuning (SFT), direct preference optimization (DPO), and inference-time expansion, the approach substantially enhances Qwen3-VL’s performance with minimal expert intervention, outperforming Gemini-3.1-Pro. Furthermore, it successfully re-annotates a large-scale professional video dataset, significantly improving the fine-grained controllability of video generation models such as Wan on complex prompts up to 400 words in length.
Existing text-to-video (T2V) diffusion models generate visually high-fidelity videos but often fail in scene construction and exhibit motion artifacts due to semantically ambiguous or logically inconsistent initial frames. To address this, we propose a semantic anchor-frame-driven factorized video generation framework that decouples T2V into three sequential stages: (1) LLM-guided semantic reasoning for coherent scene specification; (2) high-fidelity anchor frame generation via a text-to-image (T2I) model; and (3) temporal evolution modeling exclusively by a video diffusion model. This paradigm achieves the first explicit separation of scene construction and motion modeling responsibilities, significantly improving logical consistency and motion controllability. Integrated with prompt rewriting, visual anchoring sampling, and lightweight fine-tuning, our method achieves state-of-the-art performance on both T2V CompBench and VBench2, while reducing sampling steps by 70% without compromising visual quality.
Existing text-to-video models struggle to generate physically consistent dynamic content due to their reliance on implicit temporal modeling. This work proposes a dual-engine agent framework that, for the first time, leverages executable Blender code as a procedural intermediate representation. In this approach, an encoding agent generates programs describing scene composition and temporal evolution; a simulation engine executes these programs to produce deterministic spatiotemporal drafts, which are then refined by a video generation engine into photorealistic outputs. By decoupling procedural reasoning from high-fidelity rendering, the method significantly enhances controllability, interpretability, and physical consistency. Trained on a newly curated VideoCoCo-3K dataset comprising draft-instruction-target triplets, the model achieves state-of-the-art performance with scores of 0.558 on PhyGenBench and 77.88 on VBench-2.0.
Real-world multimodal video data suffers from high acquisition costs and limited diversity, hindering the training of large-scale multitask video understanding models. To address this challenge, this work proposes the first unified synthetic data generation framework capable of automatically producing unlimited, multitask-compatible multimodal video data. The approach introduces a visual question answering (VQA)-based fine-tuning strategy that replaces conventional caption- or instruction-based supervision with structured question-answer pairs to enhance the model’s visual reasoning and localization capabilities. Remarkably, models trained exclusively on this synthetic data achieve performance on par with or even surpassing fully supervised baselines on three distinct tasks—video object counting, video question answering, and video segmentation—demonstrating strong generalization and effectiveness across real-world benchmarks.
Existing generative video models struggle to achieve human-centric camera control aligned with cinematic language, often producing random camera trajectories, spatial inconsistencies, and insufficient focus on the human subject. This work proposes a human-centric camera parameterization method that formalizes cinematic composition principles into computable, human-relative camera parameters for the first time. It introduces a domain-specific language (DSL) that coordinates with a multimodal large language model to map natural language instructions and human motion into cinematic keyframe shots, followed by deterministic interpolation to generate smooth, continuous camera trajectories. Evaluated on a newly curated dataset of 34K text–motion–camera aligned samples, the approach significantly outperforms existing methods on composition-oriented metrics, enabling controllable and aesthetically cinematic human-centric video generation.