OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing video editing evaluation benchmarks, which suffer from narrow task coverage, neglect of video-specific spatiotemporal, audio, and reference dimensions, and inadequate metrics for instruction fidelity—often leading to erroneous judgments due to visual priors. To this end, we propose OmniEdit-Bench, the first comprehensive benchmark for instruction-driven video editing. It systematically decomposes multi-dimensional editing tasks under both explicit and implicit instructions and introduces a four-dimensional evaluation framework encompassing accuracy, fidelity, realism, and consistency. Crucially, an accuracy-aware penalty mechanism is incorporated to ensure instruction fidelity dominates scoring. Combining human evaluation with state-of-the-art vision-language models, OmniEdit-Bench enables fine-grained analysis. Experiments reveal that current models remain immature, underscoring the benchmark’s role as a reliable platform and directional guide for future progress.
📝 Abstract
Instruction-based video editing (IVE) is an emerging field with broad applications, yet evaluating editing models remains challenging. Existing benchmarks suffer from two major limitations: limited task coverage inherited from image editing, which overlooks video-specific dimensions, and inadequate metrics that fail to measure instruction fidelity, allowing incorrect edits to receive high scores due to strong visual priors from the original video. To address these issues, we introduce a comprehensive and structured benchmark for IVE. Our benchmark decomposes editing tasks into multiple video-specific dimensions, including spatial, temporal, audio, and reference-based editing, extending beyond conventional frame-level evaluation. It also distinguishes explicit and implicit instructions and incorporates reasoning-based scenarios to better reflect real-world requirements. Furthermore, we propose an evaluation framework that assesses editing quality from four complementary dimensions: accuracy, preservation, realism, and consistency, using both human judgments and state-of-the-art vision-language models. To emphasize instruction fidelity, we introduce an accuracy-aware penalty mechanism that conditions other scores on accuracy, preventing visually plausible but incorrect edits from receiving inflated evaluations. Extensive experiments on representative open-source and commercial models show that current IVE models remain far from satisfactory. OmniEdit-Bench provides a comprehensive and reliable testbed for evaluating instruction-based video editing and offers insights into future research directions.
Problem

Research questions and friction points this paper is trying to address.

instruction-based video editing
video editing benchmark
instruction fidelity
evaluation metrics
video-specific dimensions
Innovation

Methods, ideas, or system contributions that make the work stand out.

instruction-based video editing
video-specific evaluation
accuracy-aware penalty
structured benchmark
vision-language models