Score
Designs and builds systems that generate synthetic video clips conditioned on textual inputs, producing text-to-video or action-conditional sequences for use as augmented training data. Implements and evaluates generative augmentation pipelines to create class-conditional or label-consistent videos that balance class distributions and expand scarce classes for model training and analysis.
Text-to-video (T2V) generation remains hindered by weak text–visual alignment, poor long-term temporal coherence, and excessive computational overhead. This survey systematically traces the evolution of T2V models—from early GAN- and VAE-based approaches to contemporary diffusion-Transformer hybrids—analyzing architectural principles, pivotal advances, and drivers of paradigm shifts. It introduces the first unified framework integrating architecture design, training strategies, evaluation methodologies, and benchmark datasets; proposes a novel “perceptual alignment”-oriented evaluation paradigm; and conducts comprehensive, reproducible benchmarking of state-of-the-art models across standard datasets, exposing critical limitations of existing metrics while providing hardware specifications and hyperparameter configurations essential for replication. The study further identifies two key future directions: efficient lightweight generation and long-range spatiotemporal consistency.
Text-to-video (T2V) generation suffers from semantic and temporal inconsistency when conditioned on long, complex textual prompts. Method: We propose a zero-training, neural-symbolic feedback-based video post-processing optimization framework. It uniquely integrates symbolic reasoning—namely, event-logic modeling and object-relation verification—with neural feature alignment—specifically, cross-modal prompt-video consistency assessment—to automatically parse formal video representations and localize/correct frame-level semantic and temporal errors. Contribution/Results: Without model retraining or fine-tuning, our method improves prompt alignment by nearly 40% across multiple state-of-the-art T2V models. It significantly enhances dynamic event logicality and spatiotemporal consistency among multiple objects, establishing a new paradigm for efficient, interpretable T2V post-optimization.
To address insufficient generative image diversity—which hinders downstream classification model training—this paper introduces the “Augmented Conditioning” paradigm. Without fine-tuning pre-trained text-to-image diffusion models, it jointly conditions generation on both standard-augmented real images (e.g., cropping, color jittering) and textual prompts. This approach significantly enhances visual diversity of generated samples while preserving semantic fidelity and domain consistency. As a lightweight, inference-time, zero-shot, parameter-free strategy, it achieves state-of-the-art performance across five long-tailed and extreme few-shot (1-shot/5-shot) classification benchmarks. It improves average accuracy by 3.2% on long-tailed tasks and up to 18.7% in few-shot settings, with particularly notable gains in generalization to tail and rare classes.
Current text-to-video generation models (e.g., Sora, Gen-3) are predominantly evaluated using metrics emphasizing visual quality and motion smoothness, while neglecting temporal fidelity and text-video alignment—critical requirements for safety-critical applications. To address this gap, we propose NeuS-V, the first quantitative evaluation framework grounded in neural-symbolic formal verification. Our method comprises three key components: (1) automatic compilation of natural language prompts into temporal logic (TL) specifications; (2) symbolic modeling of videos as finite-state automata; and (3) rigorous formal verification via model checking. To support evaluation of temporal complexity, we construct the first synthetic prompt dataset explicitly designed for varying temporal intricacy. Experiments demonstrate that NeuS-V achieves over fivefold higher correlation with human judgment compared to existing metrics and, for the first time, systematically exposes severe temporal reasoning failures of state-of-the-art models under temporally complex prompts.
This work addresses the challenge of high-quality multimodal video generation and editing by proposing a unified multimodal foundation model architecture. Methodologically, it introduces variable-aspect-ratio 1080p video latent-space modeling, cross-modal alignment training across text, image, video, and audio modalities, efficient tokenization, large-scale parallel training and inference optimization, and a rigorously quality-controlled data curation strategy coupled with a novel evaluation protocol. Key contributions include the first 30-billion-parameter video generation model supporting long-horizon generation (73K tokens, i.e., 16 seconds at 16 fps), instruction-driven precise editing, user-provided image personalization, and synchronized audio-video synthesis. The model achieves state-of-the-art performance across five benchmarks: text-to-video, video personalization, video editing, video-to-audio, and text-to-audio—demonstrating substantial improvements in temporal coherence and semantic controllability.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
This study addresses the limitation of existing video generation models in rendering physically plausible interactions and state transitions. To this end, we propose a video generation framework based on a controllable interaction synthesis dataset. Specifically, we first construct a structured interaction taxonomy and leverage image editing models to generate start- and end-state anchors. Subsequently, we introduce State-Guided Sampling to achieve seamless video synthesis. Furthermore, an automated evaluation pipeline aligned with human judgment is designed to optimize data quality. Experimental results demonstrate that fine-tuning base models with our approach yields significant improvements in both the physical plausibility and visual quality of generated interactive videos.
This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.
This study addresses the challenge of maintaining cross-shot consistency in long video generation, where existing retrieval methods suffer from information redundancy or insufficient coverage. To this end, this work proposes a Complementary Retrieval-Augmented Prompting framework. By parsing screenplays to construct a visual element registry, the method introduces a novel element-aware complementary retrieval mechanism that intelligently aggregates historical reference images. This provides frozen video generators with comprehensive yet low-noise conditional constraints, enabling training-free coherent long video generation. Experimental results demonstrate that the proposed framework significantly outperforms baseline methods on multi-shot story generation tasks, substantially improving both cross-shot consistency and text controllability while exhibiting high interpretability.
This work addresses the scarcity of high-quality labeled data in 3D skeletal action recognition by proposing a conditional generative data augmentation approach constrained by action labels. Leveraging a Transformer-based encoder-decoder architecture, the method integrates a generation refinement module and a dropout mechanism to effectively balance fidelity and diversity during sequence sampling. The resulting synthetic skeletal sequences exhibit both high realism and substantial variability, consistently enhancing the performance of diverse action recognition models under both few-shot and full-data settings. Extensive experiments on the HumanAct12 and NTU-VIBE datasets demonstrate the effectiveness and generalizability of the proposed augmentation strategy.
This work addresses the challenge of poor recognition performance on rare actions in video action recognition due to long-tailed data distributions. To mitigate this issue, the authors propose the first approach that leverages text-to-video generative models for data balancing. Specifically, they construct action-semantic-guided, diverse textual prompts to synthesize videos and introduce a two-stage training strategy to alleviate domain shift between real and generated data. Remarkably, the method achieves substantial performance gains with only partial data balancing while significantly reducing computational overhead. On the UCF-LT and K100-LT benchmarks, it outperforms the current best baselines by 5.1% and 7.0%, respectively, and yields a striking 31.9% improvement on rare action categories in RareAct. Notably, it attains 79% of the full performance gain at merely 27% of the computational cost.