Score
Designs and builds curated video datasets by specifying capture and quality criteria, selecting and filtering footage (e.g., for on‑screen sound), producing temporal and semantic annotations, and partitioning the data into independent training/validation/test splits with accompanying metadata and evaluation protocols.
This work proposes a systematic data engineering methodology to address the engineering challenges of processing and managing raw video data for large-scale video foundation model training. By leveraging metadata-driven data cleaning, multi-stage filtering, and an inference-aware architecture—combined with the Lavender Data management system, μP hyperparameter transfer, and hyperspherical geometric constraint optimization—the authors efficiently construct a high-quality training set from approximately 50 million video clips. Using this curated dataset, they successfully train Summer-22B, a 22-billion-parameter video foundation model. This study provides the first empirical validation that systematic data engineering plays a dominant role in enhancing model performance, underscoring the critical importance of data quality and structural integrity in large-scale video modeling.
This work addresses the high barrier, static nature, and limited evolvability of specialized video training datasets by proposing a configurable, self-evolving video data operating system. The system enables users to issue requests via natural language and parameters, automatically optimizing queries and executing parallel retrieval of real videos alongside controllable synthesis to produce domain-specific data packages enriched with full provenance, multidimensional metadata, and reproducible notebooks. Built upon the Model-Context Protocol (MCP), it establishes a dynamic data ecosystem that supports community contributions, governance-driven continuous updates, and flexible “cooking” mechanisms. Experiments demonstrate that this approach substantially reduces dataset construction costs and enhances the training efficiency and iterative capability of multimodal large models in vertical domains.
Existing public video datasets are insufficient for advancing cinematic ultra-high-definition (UHD) text-to-video generation at 4K/8K resolutions. To address this gap, we introduce UltraVideo—the first open-source UHD-4K/8K video dataset, encompassing 100+ diverse topics and featuring per-video structured, multi-granularity captions (including 824-word detailed summaries). We design a four-stage automated curation pipeline integrating statistical filtering, large-language-model-driven purification, multimodal caption generation, and high-fidelity UHD video acquisition with precise caption alignment. Concurrently, we release UltraWan-1K and UltraWan-4K—foundation models natively optimized for 1K and 4K video generation. Notably, 22.4% of UltraVideo consists of native 8K videos, yielding substantial improvements in generation fidelity and text controllability. All data and models are publicly released under open licenses to foster reproducible research in UHD text-to-video synthesis.
To address the high cost and low efficiency of manual annotation in video object detection and segmentation, this paper proposes a lightweight, end-to-end automated annotation framework. Methodologically, it introduces the first deep integration of YOLOv8-based video tracking and SAM-based interactive segmentation, augmented by an adaptive frame-sampling strategy and a Gradio-powered visual interface, resulting in a modular, open-source, and extensible prototype system. Key contributions include: (i) a lightweight co-design of tracking and segmentation models; (ii) efficient, interactive annotation support for multi-object, long-duration videos; and (iii) rapid generation of annotated datasets enabling a closed-loop annotation–training pipeline. Experiments on multiple public video benchmarks demonstrate over 10× higher annotation throughput compared to manual labeling, strong inter-annotator consistency, and substantial performance gains for downstream detection and segmentation models trained on the generated data.
Existing video datasets suffer from coarse-grained temporal segmentation, unstructured text descriptions, and the absence of rigorous quality assessment—limiting fine-grained alignment in controllable video generation. To address these challenges, Koala-36M introduces a systematic solution: (1) constructing a large-scale dataset comprising 36 million high-quality video clips; (2) proposing the Video Training Suitability Score (VTSS), a novel multi-metric quality evaluation framework; (3) developing a high-accuracy transition detection method based on linear classification over probability distributions; and (4) generating structured, long-form textual descriptions averaging 200 words per clip, coupled with optimized fine-grained conditional embeddings. Experiments demonstrate substantial improvements in spatiotemporal text-video alignment accuracy and generation consistency. The codebase, dataset, and full preprocessing pipeline are publicly released.
This work addresses the disconnect between model design and dataset construction in existing video quality assessment (VQA) research, where efficient data selection mechanisms targeting model weaknesses are lacking. The authors propose a model-guided data selection approach that uniquely integrates failure prediction with semantic diversity. Specifically, they employ a ranking-based failure predictor to estimate sample difficulty and leverage deep semantic features to quantify content diversity, then apply a greedy algorithm to balance these two criteria for selecting high-value unlabeled videos. Using only a 5% curated subset for fine-tuning, the method improves the average SRCC from 0.651 to 0.722 and achieves top performance on the gMAD benchmark, demonstrating substantially enhanced generalization capability.
This work addresses the scarcity of large-scale, publicly available, and reproducible benchmark datasets in multimodal recommendation research. To this end, we present the first end-to-end traceable data processing pipeline built upon MovieLens-10M and MovieLens-20M, enriching them with multimodal signals—including plot summaries, movie posters, and trailer videos—and extracting textual, visual, audio, and video features using state-of-the-art encoders. The resulting datasets, M3L-10M and M3L-20M, constitute two large-scale multimodal benchmarks for movie recommendation. We provide complete mappings to original data sources, detailed feature extraction protocols, and releases in multiple formats to significantly enhance reproducibility. Both qualitative and quantitative analyses validate the effectiveness and utility of the proposed datasets, thereby advancing research in multimodal recommender systems.
为解决视频预训练数据管道封闭问题,提出VIDAFORGE,一个可执行五阶段工作流的开放研究基础设施,通过对比不同数据配方对模型性能的影响来优化视频预训练。
Existing audiovisual quality assessment datasets are limited in scale, lack diversity in content and quality degradation types, and provide only holistic scores, thereby hindering research on multimodal perception mechanisms. To address these limitations, this work proposes a crowdsourced subjective evaluation framework that transcends traditional laboratory constraints, integrating a systematic data sampling strategy with a multidimensional annotation scheme. This approach yields YT-NTU-AVQ, the largest and most diverse audiovisual quality assessment dataset to date, comprising 1,620 user-generated videos spanning a broad spectrum of semantic scenarios and quality levels. The dataset and associated platform code have been publicly released, significantly advancing the study and development of multimodal perceptual modeling.
Current research on AI-generated video detection is hindered by limited-scale datasets, outdated generative models, insufficient semantic diversity, and the absence of a systematic benchmark. To address these limitations, this work introduces AIGVDBench—the first large-scale, high-quality, and technologically representative benchmark, encompassing 31 state-of-the-art generative models and over 440,000 videos. The authors conduct more than 1,500 evaluations across 33 detectors spanning four major categories, employing multidimensional metrics and eight in-depth analyses. This comprehensive study reveals four novel findings, identifies critical performance bottlenecks and generalization patterns, and publicly releases all data and code to foster systematic progress in the field.