Score
Designs, builds, and analyzes end-to-end systems, processes, and artifacts for creating, editing, processing, and delivering video and related multimedia content, including short-form and scripted formats. Work encompasses authoring scripts and storyboards, implementing editing and processing pipelines and workflows, developing intelligent or interactive editing tools, and optimizing production processes for quality, efficiency, and distribution.
研究提出了一种自动视频编辑方法,用于生成场景预览、视频摘要和电影预告片,探讨了自动化与创意性之间的关系。
Existing procedural material generation methods merely replicate node graph structures without capturing the underlying design logic employed by experts, often yielding suboptimal results. This work proposes a process-driven generation paradigm that, for the first time, treats expert creation processes as first-class representations. By automatically analyzing tutorial videos, the approach extracts textualized process trajectories that encode design steps, parameter settings, and intent. Leveraging pretrained large language models, it constructs a ProcessSynthesizer and a Compiler to generate user-aligned trajectories and compile them into executable Blender material graphs. Expert evaluations demonstrate that the generated materials better reflect professional design strategies and require fewer edits, while a user study with 150 participants confirms significant improvements over existing systems in both output quality and editing efficiency.
In long-form narrative video editing, creators face high cognitive load when locating plot points, tracking character motivations, and reassembling distributed events across multi-hour footage; existing transcription- or embedding-based methods lack narrative understanding and fail to support creative decision-making. Method: We propose the first prompt-driven, modular video editing architecture integrating semantic indexing, temporal segmentation, guided memory compression, and cross-granularity narrative fusion to enable interpretable modeling of plot, dialogue, emotion, and context. The system replaces traditional timeline operations with natural language prompts, balancing automation efficiency with editorial control. Contribution/Results: Evaluated on 400+ videos, our method significantly improves editing efficiency while strictly preserving narrative coherence. Professional editors rated it highly in usability and creative efficacy, and user studies demonstrated strong preference over baseline approaches.
This work addresses key challenges in long-form video editing, including maintaining narrative coherence across multiple stages and enabling precise error localization and localized correction. The authors propose an open-source, multimodal multi-agent system that structures the editing process into three phases: asset preparation, editing research, and timeline execution. A novel traceable and replayable editing trajectory mechanism is introduced, allowing for accurate diagnosis of failed segments and selective re-editing without requiring full pipeline re-execution. The system integrates multi-agent collaboration, multimodal analysis, tool invocation logging, intermediate rendering, and a verifiable reward design. Experimental evaluation across 23 themes demonstrates that the approach achieves an average human rating of 3.40 out of 5, significantly outperforming CapCut-Mate and CutClaw in thematic relevance, narrative coherence, and editing fluency.
This work addresses the challenge of inefficiently adapting image-level stylization tools to frame sequences in video style transfer. We propose an automated video post-production workflow leveraging Adobe Photoshop Actions integrated with batch-processing systems. Methodologically, we systematically extend Photoshop Actions to video-scale batch rendering, establishing an end-to-end pipeline supporting frame-sequence import/export, script-driven automation, and non-destructive iterative refinement. Our contributions are threefold: (1) seamless integration of high-fidelity image-level editing capabilities with per-frame video processing; (2) pixel-accurate consistency—zero deviation—in color grading, filter application, and compositing across hundreds of frames; and (3) substantial efficiency gains in stylized editing, enabled by real-time preview support. The framework establishes a novel paradigm for lightweight, high-consistency video enhancement in creative industries.
本文提出了一种结构化协调层,通过多代理框架和FilmDSL语言,在剧本到视频生成中实现更可控和一致的长篇叙事到电影转换。
This study addresses the challenge faced by non-expert users in generating high-quality cinematic videos that balance professional storytelling with creativity. To this end, it proposes a prompt optimization framework grounded in a reusable cinematic skill library. Methodologically, the approach pioneers the evolution of cinematic skills from expert seeds, integrating resonance, incongruity, and divergent reference strategies to balance fidelity and creativity, while leveraging divergent near-miss cases to stimulate alternative ideas. Technically, it introduces fine-grained cinematic cue representations and a multi-category retrieval-augmented generation mechanism. Experimental results demonstrate that the proposed method outperforms the strongest baseline by 1.40 points on StoryEval and VBench, significantly surpassing seed skills, and establishes a comprehensive four-dimensional evaluation framework.
This study addresses the cumbersome processes of localization, segmentation, and prompt construction in multi-shot generative video editing by proposing an interaction paradigm grounded in multi-level structural parsing. The method transforms videos into malleable hierarchical structures, enabling users to modify elements within a task-centric workspace while AI agents automatically handle intent translation and change propagation. Based on this approach, an interactive system is developed to support both rapid prototyping and end-to-end post-production workflows. User studies and expert evaluations demonstrate that the system significantly enhances efficiency in video comprehension, intent expression, and solution exploration, thereby establishing an effective new paradigm for generative video editing.
本文提出FRAMEWORKERS框架,通过多代理动态协调解决AI视频制作中的任务管理和资产调度问题,提高视频质量和任务覆盖范围。
This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.