Score
Designs, produces, edits, and packages short‑duration video content, covering scripting, storyboarding, shot composition, pacing, visual and audio editing, captioning/subtitles, and export formats. Builds versions tuned to platform constraints and analyzes engagement and performance metrics to iterate on length, format, and creative choices.
研究提出了一种自动视频编辑方法,用于生成场景预览、视频摘要和电影预告片,探讨了自动化与创意性之间的关系。
研究通过专业编辑对AI生成的电影广告评价,提出六维评估框架,指导AI生成广告的编辑意识、人工及自动评估。
Existing automatic video editing methods over-rely on ASR transcripts while neglecting visual context, resulting in narratively incoherent outputs. To address this, we propose HIVE, the first end-to-end video editing framework integrating multimodal narrative understanding. HIVE decomposes editing into three stages—highlight detection, head/tail selection, and redundancy removal—guided by character extraction, dialogue analysis, and narrative summarization. It synergistically fuses multimodal large language models, scene-level segmentation, ASR transcripts, and visual context modeling to emulate human editorial reasoning. Evaluated on our newly constructed dataset DramaAD, HIVE achieves significant improvements in narrative coherence and readability for both general and advertisement-oriented short-video generation. Quantitatively, it substantially narrows the quality gap between automated editing and professional human editing, demonstrating superior alignment with human perceptual and narrative expectations.
In long-form narrative video editing, creators face high cognitive load when locating plot points, tracking character motivations, and reassembling distributed events across multi-hour footage; existing transcription- or embedding-based methods lack narrative understanding and fail to support creative decision-making. Method: We propose the first prompt-driven, modular video editing architecture integrating semantic indexing, temporal segmentation, guided memory compression, and cross-granularity narrative fusion to enable interpretable modeling of plot, dialogue, emotion, and context. The system replaces traditional timeline operations with natural language prompts, balancing automation efficiency with editorial control. Contribution/Results: Evaluated on 400+ videos, our method significantly improves editing efficiency while strictly preserving narrative coherence. Professional editors rated it highly in usability and creative efficacy, and user studies demonstrated strong preference over baseline approaches.
Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.
本文提出了一种结构化协调层,通过多代理框架和FilmDSL语言,在剧本到视频生成中实现更可控和一致的长篇叙事到电影转换。
本文提出FRAMEWORKERS框架,通过多代理动态协调解决AI视频制作中的任务管理和资产调度问题,提高视频质量和任务覆盖范围。
This study addresses the challenge faced by non-expert users in generating high-quality cinematic videos that balance professional storytelling with creativity. To this end, it proposes a prompt optimization framework grounded in a reusable cinematic skill library. Methodologically, the approach pioneers the evolution of cinematic skills from expert seeds, integrating resonance, incongruity, and divergent reference strategies to balance fidelity and creativity, while leveraging divergent near-miss cases to stimulate alternative ideas. Technically, it introduces fine-grained cinematic cue representations and a multi-category retrieval-augmented generation mechanism. Experimental results demonstrate that the proposed method outperforms the strongest baseline by 1.40 points on StoryEval and VBench, significantly surpassing seed skills, and establishes a comprehensive four-dimensional evaluation framework.
Existing automated video editing systems struggle to simultaneously support diverse editing tasks and maintain coherent narrative structures in long-form videos. This work proposes a unified multi-agent framework that automates shot creation through a dedicated shot-planning agent and dynamically orchestrates over thirty specialized editing agents by integrating intent parsing with text-to-gradient map optimization. By combining cross-modal retrieval and large-scale tool integration, the method significantly outperforms current approaches on the VideoEdit benchmark and public datasets, achieving orchestration success rates of 87–95%, reducing API costs by 60%, and attaining human evaluation scores only 4% below those of professional human editors.
Existing text- or prompt-driven video generation methods struggle to meet the demands of short-form dramas, which require rapid shot transitions, dialogue-driven focus shifts, and cinematic visual composition. To address this, we propose a geometry-guided framework for short drama generation that decouples static visual structure from dynamic narrative conditions, leveraging depth and pose priors to guide initial frame synthesis and image-to-video generation. We introduce DramaBoard, the first structured storyboard dataset for short dramas, and design a constrained training mechanism incorporating text–visual alignment rewards, schema-constrained supervised fine-tuning, and GRPO-based reinforcement learning. Experiments demonstrate that our approach significantly outperforms existing baselines in terms of fidelity, temporal consistency, and controllability. We publicly release our code and the DramaBoard evaluation benchmark.