🤖 AI Summary
This study addresses the fragmentation between video generation and editing in existing datasets, which fails to reflect the iterative creative workflows of AI-native video production. To this end, we propose Sora100K, a dataset that models video creation as a structured trajectory paradigm comprising generation root nodes, editing edges, and multi-turn interaction logs. A rigorous pipeline is constructed to reconstruct complete provenance relationships and intermediate states from source to edited videos. Methodologically, vision-language models are employed for semantic and operational annotation, while lightweight adaptation using the LTX-2 model evaluates supervisory value. Experiments demonstrate that this paradigm significantly enhances visual quality, multi-shot generation capabilities, and cross-shot consistency, validating the potential of trajectory data in supporting complex editing instructions.
📝 Abstract
AI-Native video creation is shifting from isolated video clips toward iterative video creation workflows. However, existing datasets remain largely video clips, representing video generation and editing as separate tasks rather than connected stages of a video creation workflow. In this paper, we introduce Sora100K, a dataset that represents the AI-Native video creation workflow as a structured video creation trajectory. Specifically, we first identify video creation trajectories and decompose them into three subsets according to their structural roles: text-to-video generation records as roots, single-turn video editing records as editing edges, and multi-turn video editing records as complete trajectories. Then, we use a VLM to assign semantic annotations for generation roots and editing-operation annotations for editing edges. A strict construction pipeline further reconstructs source-to-edit lineage, editing order, and intermediate video states while ensuring data quality. Finally, we perform lightweight adaptation on LTX-2 models to assess the supervision value of Sora100K. The results show improvements in visual quality, multi-shot generation, and cross-shot consistency, while successive-turn evaluation reveals that following multi-turn editing instructions remains challenging. Sora100K establishes a new data foundation for AI-Native video creation beyond isolated video clips and toward structured video creation trajectory. The dataset and supplementary materials are publicly available at https://huggingface.co/datasets/ysicong/Sora100K.