video production

Designs, builds, and analyzes end-to-end systems, processes, and artifacts for creating, editing, processing, and delivering video and related multimedia content, including short-form and scripted formats. Work encompasses authoring scripts and storyboards, implementing editing and processing pipelines and workflows, developing intelligent or interactive editing tools, and optimizing production processes for quality, efficiency, and distribution.

videoproduction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.38
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$197K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing procedural material generation methods merely replicate node graph structures without capturing the underlying design logic employed by experts, often yielding suboptimal results. This work proposes a process-driven generation paradigm that, for the first time, treats expert creation processes as first-class representations. By automatically analyzing tutorial videos, the approach extracts textualized process trajectories that encode design steps, parameter settings, and intent. Leveraging pretrained large language models, it constructs a ProcessSynthesizer and a Compiler to generate user-aligned trajectories and compile them into executable Blender material graphs. Expert evaluations demonstrate that the generated materials better reflect professional design strategies and require fewer edits, while a user study with 150 participants confirms significant improvements over existing systems in both output quality and editing efficiency.

design intentexpert demonstrationsmaterial authoring

Prompt-Driven Agentic Video Editing System: Autonomous Comprehension of Long-Form, Story-Driven Media

Sep 20, 2025
ZD
Zihan Ding
🏛️ University of British Columbia | University of Cambridge | University of Bristol | Memories.AI

In long-form narrative video editing, creators face high cognitive load when locating plot points, tracking character motivations, and reassembling distributed events across multi-hour footage; existing transcription- or embedding-based methods lack narrative understanding and fail to support creative decision-making. Method: We propose the first prompt-driven, modular video editing architecture integrating semantic indexing, temporal segmentation, guided memory compression, and cross-granularity narrative fusion to enable interpretable modeling of plot, dialogue, emotion, and context. The system replaces traditional timeline operations with natural language prompts, balancing automation efficiency with editorial control. Contribution/Results: Evaluated on 400+ videos, our method significantly improves editing efficiency while strictly preserving narrative coherence. Professional editors rated it highly in usability and creative efficacy, and user studies demonstrated strong preference over baseline approaches.

Creators need systems that preserve narrative coherence while allowing prompt-driven editing controlEditing long narrative videos requires overcoming cognitive demands of storyboarding and sequencingExisting methods fail to track characters and connect dispersed events in creative workflows

This work addresses key challenges in long-form video editing, including maintaining narrative coherence across multiple stages and enabling precise error localization and localized correction. The authors propose an open-source, multimodal multi-agent system that structures the editing process into three phases: asset preparation, editing research, and timeline execution. A novel traceable and replayable editing trajectory mechanism is introduced, allowing for accurate diagnosis of failed segments and selective re-editing without requiring full pipeline re-execution. The system integrates multi-agent collaboration, multimodal analysis, tool invocation logging, intermediate rendering, and a verifiable reward design. Experimental evaluation across 23 themes demonstrates that the approach achieves an average human rating of 3.40 out of 5, significantly outperforming CapCut-Mate and CutClaw in thematic relevance, narrative coherence, and editing fluency.

failure diagnosislong-form video editingmulti-agent system

Photoshop Batch Rendering Using Actions for Stylistic Video Editing

May 02, 2025
TD
Tessa De La Fuente
🏛️ Texas A&M University

This work addresses the challenge of inefficiently adapting image-level stylization tools to frame sequences in video style transfer. We propose an automated video post-production workflow leveraging Adobe Photoshop Actions integrated with batch-processing systems. Methodologically, we systematically extend Photoshop Actions to video-scale batch rendering, establishing an end-to-end pipeline supporting frame-sequence import/export, script-driven automation, and non-destructive iterative refinement. Our contributions are threefold: (1) seamless integration of high-fidelity image-level editing capabilities with per-frame video processing; (2) pixel-accurate consistency—zero deviation—in color grading, filter application, and compositing across hundreds of frames; and (3) substantial efficiency gains in stylized editing, enabled by real-time preview support. The framework establishes a novel paradigm for lightweight, high-consistency video enhancement in creative industries.

Automated visual edits across multiple imagesEfficient workflow for creative image/video editingOptimizing productivity with uniform editing results

Latest Papers

What's happening recently
View more

This study addresses the challenge faced by non-expert users in generating high-quality cinematic videos that balance professional storytelling with creativity. To this end, it proposes a prompt optimization framework grounded in a reusable cinematic skill library. Methodologically, the approach pioneers the evolution of cinematic skills from expert seeds, integrating resonance, incongruity, and divergent reference strategies to balance fidelity and creativity, while leveraging divergent near-miss cases to stimulate alternative ideas. Technically, it introduces fine-grained cinematic cue representations and a multi-category retrieval-augmented generation mechanism. Experimental results demonstrate that the proposed method outperforms the strongest baseline by 1.40 points on StoryEval and VBench, significantly surpassing seed skills, and establishes a comprehensive four-dimensional evaluation framework.

cinematic qualitycreativityprompt engineering

This study addresses the cumbersome processes of localization, segmentation, and prompt construction in multi-shot generative video editing by proposing an interaction paradigm grounded in multi-level structural parsing. The method transforms videos into malleable hierarchical structures, enabling users to modify elements within a task-centric workspace while AI agents automatically handle intent translation and change propagation. Based on this approach, an interactive system is developed to support both rapid prototyping and end-to-end post-production workflows. User studies and expert evaluations demonstrate that the system significantly enhances efficiency in video comprehension, intent expression, and solution exploration, thereby establishing an effective new paradigm for generative video editing.

generative video editinginteractive editinglong-form video

This work addresses the challenge in generative video editing where object-level geometric manipulations—such as translation, rotation, scaling, duplication, or deletion—often fail to consistently update secondary visual effects like shadows and reflections. To this end, the authors propose GIVE, a unified framework that models pre- and post-edit 3D geometric changes through a consistent object state representation. GIVE employs a dual geometric stream composed of depth and orientation boxes to generate compact, temporally aligned editing instructions. The framework leverages a scalable, procedural synthetic data pipeline built upon a graphics engine for supervised training. GIVE is the first to support diverse geometric editing operations within a single architecture while explicitly modeling 3D state transitions, thereby ensuring consistency in secondary effects, high visual fidelity, temporal coherence, and strong generalization to real-world videos.

3D object manipulationgeometric editingsecondary effects

Hot Scholars

RB

Roberto Balestri

Ph. D. Student
Artificial IntelligenceNarrative AnalysisMedia StudiesCinema and Television
ZZ

Zhaoxiang Zhang

Institute of Automation, Chinese Academy of Sciences
Computer VisionPattern RecognitionBiologically-inspired Learning
JS

Josef Sivic

Czech Technical University, CIIRC, ELLIS Unit Prague
computer visionmachine learning
FC

Fabian Caba Heilbron

Research Assistant, King Abdullah University of Science and Technology
Computer Vision