🤖 AI Summary
This study addresses the cumbersome processes of localization, segmentation, and prompt construction in multi-shot generative video editing by proposing an interaction paradigm grounded in multi-level structural parsing. The method transforms videos into malleable hierarchical structures, enabling users to modify elements within a task-centric workspace while AI agents automatically handle intent translation and change propagation. Based on this approach, an interactive system is developed to support both rapid prototyping and end-to-end post-production workflows. User studies and expert evaluations demonstrate that the system significantly enhances efficiency in video comprehension, intent expression, and solution exploration, thereby establishing an effective new paradigm for generative video editing.
📝 Abstract
Recent generative video editing models enable video content modification (e.g., changing a character) but target short clips. Extending them to full multi-shot videos requires tedious work to locate relevant content across shots, segment it into clips, craft context-aware editing prompts for each clip, and repeatedly articulate complex editing intent. To address this, we explore an interaction paradigm for editing through underlying video structures (e.g., scripts, scenes, characters, shots, and their relationships). We present CraftTrace, an interactive prototype that transforms a video into a malleable, multilevel structure for generative editing. Users work in task-centric workspaces to modify elements or reshape relationships, while an AI agent translates and propagates changes across the video. A user study and expert review show that this structure helps users understand videos, formulate and refine editing intent, and explore alternatives, supporting rapid prototyping during early-stage exploration and full video post-production.