short video

Designs, produces, edits, and packages short‑duration video content, covering scripting, storyboarding, shot composition, pacing, visual and audio editing, captioning/subtitles, and export formats. Builds versions tuned to platform constraints and analyzes engagement and performance metrics to iterate on length, format, and creative choices.

shortvideo

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.59
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Jul 03, 2025
XW
Xiangfeng Wang
🏛️ University of Science and Technology of China | ByteDance China

Existing automatic video editing methods over-rely on ASR transcripts while neglecting visual context, resulting in narratively incoherent outputs. To address this, we propose HIVE, the first end-to-end video editing framework integrating multimodal narrative understanding. HIVE decomposes editing into three stages—highlight detection, head/tail selection, and redundancy removal—guided by character extraction, dialogue analysis, and narrative summarization. It synergistically fuses multimodal large language models, scene-level segmentation, ASR transcripts, and visual context modeling to emulate human editorial reasoning. Evaluated on our newly constructed dataset DramaAD, HIVE achieves significant improvements in narrative coherence and readability for both general and advertisement-oriented short-video generation. Quantitatively, it substantially narrows the quality gap between automated editing and professional human editing, demonstrating superior alignment with human perceptual and narrative expectations.

Bridge quality gap between automatic and human editingCondense long videos into engaging short clipsImprove coherence using multimodal narrative understanding

Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media

Sep 20, 2025
ZD
Zihan Ding
🏛️ University of British Columbia | University of Cambridge | University of Bristol | Memories.AI

In long-form narrative video editing, creators face high cognitive load when locating plot points, tracking character motivations, and reassembling distributed events across multi-hour footage; existing transcription- or embedding-based methods lack narrative understanding and fail to support creative decision-making. Method: We propose the first prompt-driven, modular video editing architecture integrating semantic indexing, temporal segmentation, guided memory compression, and cross-granularity narrative fusion to enable interpretable modeling of plot, dialogue, emotion, and context. The system replaces traditional timeline operations with natural language prompts, balancing automation efficiency with editorial control. Contribution/Results: Evaluated on 400+ videos, our method significantly improves editing efficiency while strictly preserving narrative coherence. Professional editors rated it highly in usability and creative efficacy, and user studies demonstrated strong preference over baseline approaches.

Creators need systems that preserve narrative coherence while allowing prompt-driven editing controlEditing long narrative videos requires overcoming cognitive demands of storyboarding and sequencingExisting methods fail to track characters and connect dispersed events in creative workflows

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

Aug 19, 2024
LH
Liu He
🏛️ Purdue University | Baidu

Current text-to-video models suffer from significant deficiencies in physical plausibility, photorealistic lighting, camera motion, and temporal coherence, limiting their applicability to cinematic-grade synthesis. To address this, we propose the first multi-agent VLM framework tailored for high-fidelity 3D video generation, featuring decoupled Director, Programmer, and Reviewer agents. Our method decomposes the synthesis task, automatically generates Blender scripting code, and performs iterative optimization guided by vision-language feedback—enabling end-to-end, interpretable, and editable video generation. Deeply integrating cinematographic knowledge with a closed-loop 3D rendering pipeline, it produces high-fidelity videos fully aligned with textual prompts—without manual intervention. Experiments demonstrate superior performance over leading commercial models across five video quality and instruction-following metrics. User studies further confirm substantial improvements: +28.6% in physical plausibility, +31.2% in temporal consistency, and higher overall quality scores.

Address improper motion and consistency in text-to-video generationAutomate synthetic video creation via VLM agent collaborationReduce manual CGI editing in film industry workflows

Latest Papers

What's happening recently
View more

This study addresses the challenge faced by non-expert users in generating high-quality cinematic videos that balance professional storytelling with creativity. To this end, it proposes a prompt optimization framework grounded in a reusable cinematic skill library. Methodologically, the approach pioneers the evolution of cinematic skills from expert seeds, integrating resonance, incongruity, and divergent reference strategies to balance fidelity and creativity, while leveraging divergent near-miss cases to stimulate alternative ideas. Technically, it introduces fine-grained cinematic cue representations and a multi-category retrieval-augmented generation mechanism. Experimental results demonstrate that the proposed method outperforms the strongest baseline by 1.40 points on StoryEval and VBench, significantly surpassing seed skills, and establishes a comprehensive four-dimensional evaluation framework.

cinematic qualitycreativityprompt engineering

Existing automated video editing systems struggle to simultaneously support diverse editing tasks and maintain coherent narrative structures in long-form videos. This work proposes a unified multi-agent framework that automates shot creation through a dedicated shot-planning agent and dynamically orchestrates over thirty specialized editing agents by integrating intent parsing with text-to-gradient map optimization. By combining cross-modal retrieval and large-scale tool integration, the method significantly outperforms current approaches on the VideoEdit benchmark and public datasets, achieving orchestration success rates of 87–95%, reducing API costs by 60%, and attaining human evaluation scores only 4% below those of professional human editors.

automated video processingcoherent narrativelong-video understanding

Existing text- or prompt-driven video generation methods struggle to meet the demands of short-form dramas, which require rapid shot transitions, dialogue-driven focus shifts, and cinematic visual composition. To address this, we propose a geometry-guided framework for short drama generation that decouples static visual structure from dynamic narrative conditions, leveraging depth and pose priors to guide initial frame synthesis and image-to-video generation. We introduce DramaBoard, the first structured storyboard dataset for short dramas, and design a constrained training mechanism incorporating text–visual alignment rewards, schema-constrained supervised fine-tuning, and GRPO-based reinforcement learning. Experiments demonstrate that our approach significantly outperforms existing baselines in terms of fidelity, temporal consistency, and controllability. We publicly release our code and the DramaBoard evaluation benchmark.

cinematographic groundingmulti-shot videoplot-to-video