🤖 AI Summary
This work addresses the limitation of existing video editing evaluation benchmarks, which overlook the diversity of editing intents and struggle to assess the effectiveness of different narrative strategies under message-driven conditions. To bridge this gap, we introduce MEDit-Bench—the first benchmark specifically designed for message-driven narrative video editing—comprising long-form videos, multiple editing instructions per video, and corresponding professional edits. We propose a temporally aligned automatic evaluation protocol and conduct a systematic analysis leveraging multimodal large language models, reinforcement-fine-tuned baselines, and human perceptual studies to investigate how message ambiguity and contextual richness impact model performance. Experiments show that state-of-the-art models approach human-level alignment under lenient temporal thresholds but remain substantially inferior under stricter criteria. Human evaluations confirm that professional edits significantly outperform model-generated outputs and reveal positional bias in LLM-based assessments, while also demonstrating that message difficulty effectively stratifies model performance.
📝 Abstract
Video editing is fundamentally message-driven: even from the same source footage, the selected shots change depending on the narrative the editor wishes to convey. Benchmarks for a closely related task, video summarization, reduce editorial intent to a single, message-agnostic notion of saliency and thus do not account for this diversity. For evaluating message-driven video editing, we present \textbf{MEDit-Bench}, a dataset and benchmark, which pairs long-form videos with multiple editing messages and multiple professionally produced edits per message, demonstrating that different messages yield substantially different edits from the same source. We define an automatic evaluation protocol based on temporal alignment metrics, and find that an LLM-as-a-judge preference, a natural proxy for narrative quality, is unreliable for this task due to severe position bias. We additionally annotate each message with ambiguity and contextfulness scores, and show that both dimensions negatively correlate with model performance, establishing message difficulty as a meaningful stratification factor. Experiments with state-of-the-art MLLMs and reinforcement fine-tuned baselines show that while strong models approach human temporal alignment at lenient thresholds, all models fall behind humans at stricter criteria. A human perceptual study further confirms a large quality gap, with professional human edits remaining consistently preferred over model outputs.