Bridging Text and Video Generation: A Survey

📅 2025-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Text-to-video (T2V) generation remains hindered by weak text–visual alignment, poor long-term temporal coherence, and excessive computational overhead. This survey systematically traces the evolution of T2V models—from early GAN- and VAE-based approaches to contemporary diffusion-Transformer hybrids—analyzing architectural principles, pivotal advances, and drivers of paradigm shifts. It introduces the first unified framework integrating architecture design, training strategies, evaluation methodologies, and benchmark datasets; proposes a novel “perceptual alignment”-oriented evaluation paradigm; and conducts comprehensive, reproducible benchmarking of state-of-the-art models across standard datasets, exposing critical limitations of existing metrics while providing hardware specifications and hyperparameter configurations essential for replication. The study further identifies two key future directions: efficient lightweight generation and long-range spatiotemporal consistency.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: GenerationMachine Learning: Deep Neural Architectures and Foundation Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating coherent visual content from natural language prompts. From its inception, the field has advanced from adversarial models to diffusion-based models, yielding higher-fidelity, temporally consistent outputs. Yet challenges persist, such as alignment, long-range coherence, and computational efficiency. Addressing this evolving landscape, we present a comprehensive survey of text-to-video generative models, tracing their development from early GANs and VAEs to hybrid Diffusion-Transformer (DiT) architectures, detailing how these models work, what limitations they addressed in their predecessors, and why shifts toward new architectural paradigms were necessary to overcome challenges in quality, coherence, and control. We provide a systematic account of the datasets, which the surveyed text-to-video models were trained and evaluated on, and, to support reproducibility and assess the accessibility of training such models, we detail their training configurations, including their hardware specifications, GPU counts, batch sizes, learning rates, optimizers, epochs, and other key hyperparameters. Further, we outline the evaluation metrics commonly used for evaluating such models and present their performance across standard benchmarks, while also discussing the limitations of these metrics and the emerging shift toward more holistic, perception-aligned evaluation strategies. Finally, drawing from our analysis, we outline the current open challenges and propose a few promising future directions, laying out a perspective for future researchers to explore and build upon in advancing T2V research and applications.
Problem

Research questions and friction points this paper is trying to address.

Surveying text-to-video generation models from early GANs to diffusion architectures
Analyzing training datasets, configurations and evaluation metrics for T2V models
Identifying current challenges and proposing future directions for video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Survey of text-to-video generative models evolution
Comparison of GANs VAEs to Diffusion-Transformer architectures
Analysis of training configurations and evaluation metrics
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Nilay Kumar
Nilay Kumar
Department of Computational Intelligence, SRM Institute of Science and Technology, KTR
P
Priyansh Bhandari
Department of Computational Intelligence, SRM Institute of Science and Technology, KTR
G
G. Maragatham
Department of Computational Intelligence, SRM Institute of Science and Technology, KTR