text-conditioned video augmentation

Designs and builds systems that generate synthetic video clips conditioned on textual inputs, producing text-to-video or action-conditional sequences for use as augmented training data. Implements and evaluates generative augmentation pipelines to create class-conditional or label-consistent videos that balance class distributions and expand scarce classes for model training and analysis.

text-conditionedvideoaugmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback

Apr 24, 2025
MC
Minkyu Choi
🏛️ The University of Texas at Austin

Text-to-video (T2V) generation suffers from semantic and temporal inconsistency when conditioned on long, complex textual prompts. Method: We propose a zero-training, neural-symbolic feedback-based video post-processing optimization framework. It uniquely integrates symbolic reasoning—namely, event-logic modeling and object-relation verification—with neural feature alignment—specifically, cross-modal prompt-video consistency assessment—to automatically parse formal video representations and localize/correct frame-level semantic and temporal errors. Contribution/Results: Without model retraining or fine-tuning, our method improves prompt alignment by nearly 40% across multiple state-of-the-art T2V models. It significantly enhances dynamic event logicality and spatiotemporal consistency among multiple objects, establishing a new paradigm for efficient, interpretable T2V post-optimization.

Enhancing temporal alignment in complex video promptsImproving semantic consistency in text-to-video generationReducing computational costs for video refinement

Augmented Conditioning Is Enough For Effective Training Image Generation

Feb 06, 2025
JC
Jiahui Chen
🏛️ UT Austin | McGill University

To address insufficient generative image diversity—which hinders downstream classification model training—this paper introduces the “Augmented Conditioning” paradigm. Without fine-tuning pre-trained text-to-image diffusion models, it jointly conditions generation on both standard-augmented real images (e.g., cropping, color jittering) and textual prompts. This approach significantly enhances visual diversity of generated samples while preserving semantic fidelity and domain consistency. As a lightweight, inference-time, zero-shot, parameter-free strategy, it achieves state-of-the-art performance across five long-tailed and extreme few-shot (1-shot/5-shot) classification benchmarks. It improves average accuracy by 3.2% on long-tailed tasks and up to 18.7% in few-shot settings, with particularly notable gains in generalization to tail and rare classes.

Augmenting conditioning for better classification benchmarksEnhancing diversity in text-to-image generation modelsImproving synthetic image effectiveness for training

Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification

Nov 22, 2024
SP
S. P. Sharan
🏛️ The University of Texas at Austin

Current text-to-video generation models (e.g., Sora, Gen-3) are predominantly evaluated using metrics emphasizing visual quality and motion smoothness, while neglecting temporal fidelity and text-video alignment—critical requirements for safety-critical applications. To address this gap, we propose NeuS-V, the first quantitative evaluation framework grounded in neural-symbolic formal verification. Our method comprises three key components: (1) automatic compilation of natural language prompts into temporal logic (TL) specifications; (2) symbolic modeling of videos as finite-state automata; and (3) rigorous formal verification via model checking. To support evaluation of temporal complexity, we construct the first synthetic prompt dataset explicitly designed for varying temporal intricacy. Experiments demonstrate that NeuS-V achieves over fivefold higher correlation with human judgment compared to existing metrics and, for the first time, systematically exposes severe temporal reasoning failures of state-of-the-art models under temporally complex prompts.

Addressing gaps in current video generation evaluation metricsAssessing text-to-video alignment in synthetic video modelsEvaluating temporal fidelity using neuro-symbolic formal verification

This work addresses the challenge of high-quality multimodal video generation and editing by proposing a unified multimodal foundation model architecture. Methodologically, it introduces variable-aspect-ratio 1080p video latent-space modeling, cross-modal alignment training across text, image, video, and audio modalities, efficient tokenization, large-scale parallel training and inference optimization, and a rigorously quality-controlled data curation strategy coupled with a novel evaluation protocol. Key contributions include the first 30-billion-parameter video generation model supporting long-horizon generation (73K tokens, i.e., 16 seconds at 16 fps), instruction-driven precise editing, user-provided image personalization, and synchronized audio-video synthesis. The model achieves state-of-the-art performance across five benchmarks: text-to-video, video personalization, video editing, video-to-audio, and text-to-audio—demonstrating substantial improvements in temporal coherence and semantic controllability.

Develop high-quality 1080p HD video generation.Enable precise instruction-based video editing.Generate personalized videos using user images.

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

Latest Papers

What's happening recently
View more

This study addresses the limitation of existing video generation models in rendering physically plausible interactions and state transitions. To this end, we propose a video generation framework based on a controllable interaction synthesis dataset. Specifically, we first construct a structured interaction taxonomy and leverage image editing models to generate start- and end-state anchors. Subsequently, we introduce State-Guided Sampling to achieve seamless video synthesis. Furthermore, an automated evaluation pipeline aligned with human judgment is designed to optimize data quality. Experimental results demonstrate that fine-tuning base models with our approach yields significant improvements in both the physical plausibility and visual quality of generated interactive videos.

physical interactionroboticsstate transition

This study investigates the impact of training data distribution and caption quality on the performance of text-to-video generation models, addressing a critical gap in the field’s data-centric understanding. To this end, we introduce Moving Alphabet, a highly controllable synthetic data platform that programmatically generates videos of moving letters with precise annotations, enabling systematic ablation studies. Our work presents the first application of such controllable synthetic data to text-to-video generation, revealing that balanced data distributions and diverse video durations substantially enhance model generalization. We further demonstrate that caption quality directly affects both training efficiency and generation fidelity. While high-quality fine-tuning can partially mitigate the limitations imposed by low-quality pretraining data, it cannot fully compensate for these deficiencies.

caption qualitydata distributiontext-to-video generation

This study addresses the challenge of maintaining cross-shot consistency in long video generation, where existing retrieval methods suffer from information redundancy or insufficient coverage. To this end, this work proposes a Complementary Retrieval-Augmented Prompting framework. By parsing screenplays to construct a visual element registry, the method introduces a novel element-aware complementary retrieval mechanism that intelligently aggregates historical reference images. This provides frozen video generators with comprehensive yet low-noise conditional constraints, enabling training-free coherent long video generation. Experimental results demonstrate that the proposed framework significantly outperforms baseline methods on multi-shot story generation tasks, substantially improving both cross-shot consistency and text controllability while exhibiting high interpretability.

Cross-shot consistencyInformational mismatchLong-form video generation

This work addresses the scarcity of high-quality labeled data in 3D skeletal action recognition by proposing a conditional generative data augmentation approach constrained by action labels. Leveraging a Transformer-based encoder-decoder architecture, the method integrates a generation refinement module and a dropout mechanism to effectively balance fidelity and diversity during sequence sampling. The resulting synthetic skeletal sequences exhibit both high realism and substantial variability, consistently enhancing the performance of diverse action recognition models under both few-shot and full-data settings. Extensive experiments on the HumanAct12 and NTU-VIBE datasets demonstrate the effectiveness and generalizability of the proposed augmentation strategy.

3D skeleton datadata augmentationhuman behaviour understanding

This work addresses the challenge of poor recognition performance on rare actions in video action recognition due to long-tailed data distributions. To mitigate this issue, the authors propose the first approach that leverages text-to-video generative models for data balancing. Specifically, they construct action-semantic-guided, diverse textual prompts to synthesize videos and introduce a two-stage training strategy to alleviate domain shift between real and generated data. Remarkably, the method achieves substantial performance gains with only partial data balancing while significantly reducing computational overhead. On the UCF-LT and K100-LT benchmarks, it outperforms the current best baselines by 5.1% and 7.0%, respectively, and yields a striking 31.9% improvement on rare action categories in RareAct. Notably, it attains 79% of the full performance gain at merely 27% of the computational cost.

data imbalancelong-tailedrare actions

Hot Scholars

XL

Xiu Li

Bytedance Seed
Computer VisionComputer Graphics3D Vision
YY

Yifan Yang

Senior Research SDE, Microsoft Research Asia
Multi-modalityComputer VisionMachine LearningArtificial Intelligence
YS

Yumeng Shi

Nanyang Technology University
efficient machine learningmultimodalityvideo understanding
SW

Shuo Wang

School of System Science, Beijing Normal University & D-ITET, ETH Zürich
Graph neural networksComputer VisionAir quality