🤖 AI Summary
Existing public video datasets are insufficient for advancing cinematic ultra-high-definition (UHD) text-to-video generation at 4K/8K resolutions. To address this gap, we introduce UltraVideo—the first open-source UHD-4K/8K video dataset, encompassing 100+ diverse topics and featuring per-video structured, multi-granularity captions (including 824-word detailed summaries). We design a four-stage automated curation pipeline integrating statistical filtering, large-language-model-driven purification, multimodal caption generation, and high-fidelity UHD video acquisition with precise caption alignment. Concurrently, we release UltraWan-1K and UltraWan-4K—foundation models natively optimized for 1K and 4K video generation. Notably, 22.4% of UltraVideo consists of native 8K videos, yielding substantial improvements in generation fidelity and text controllability. All data and models are publicly released under open licenses to foster reproducible research in UHD text-to-video synthesis.
📝 Abstract
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements for high-quality video generation models. For example, the generation of movie-level Ultra-High Definition (UHD) videos and the creation of 4K short video content. However, the existing public datasets cannot support related research and applications. In this paper, we first propose a high-quality open-sourced UHD-4K (22.4% of which are 8K) text-to-video dataset named UltraVideo, which contains a wide range of topics (more than 100 kinds), and each video has 9 structured captions with one summarized caption (average of 824 words). Specifically, we carefully design a highly automated curation process with four stages to obtain the final high-quality dataset: extit{i)} collection of diverse and high-quality video clips. extit{ii)} statistical data filtering. extit{iii)} model-based data purification. extit{iv)} generation of comprehensive, structured captions. In addition, we expand Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of our data curation.We believe that this work can make a significant contribution to future research on UHD video generation. UltraVideo dataset and UltraWan models are available at https://xzc-zju.github.io/projects/UltraVideo.