VideoX-Qwen: Data-Centric Instruction-Based Video Editing

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出VideoX-Qwen框架,通过构建大规模数据集和统一训练模型解决基于指令的视频编辑问题,提高了编辑质量和内容保真度。
📝 Abstract
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
Problem

Research questions and friction points this paper is trying to address.

video editing
instruction-driven
large-scale paired supervision
content preservation
temporal continuity
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction-based video editing
data-construction and model-training framework
multimodal semantic conditioning
dense source-video latent guidance
progressive image-video training
🔎 Similar Papers
No similar papers found.
Jiahang Li
Jiahang Li
School of Intelligence Science and Technology, Nanjing University
D
Dingbao Shao
School of Intelligence Science and Technology, Nanjing University
X
Xinyu Chen
School of Intelligence Science and Technology, Nanjing University
Song Wu
Song Wu
Southwest University
Computer VisionMachine LearningDeep learningMultimedia
Jiang Lin
Jiang Lin
StarsMicroSystem
computer architecturememory systemsoperating systems
D
Duo Li
School of Intelligence Science and Technology, Nanjing University
Yuhang Liu
Yuhang Liu
The University of Adelaide
Representation LearningLLMsLatent Variable ModelsResponsible AI
J
Jiaxin Hu
School of Intelligence Science and Technology, Nanjing University
S
Shengrong Gu
School of Intelligence Science and Technology, Nanjing University
Y
Ying Tai
School of Intelligence Science and Technology, Nanjing University
Z
Zili Yi
School of Intelligence Science and Technology, Nanjing University