FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry

πŸ“… 2026-07-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limitations of existing video editing datasets, which rely on manual annotations and suffer from large synthesis errors and insufficient task diversity, hindering models’ ability to internalize language-guided spatial localization and editing. The authors propose a mask-free, online method for generating video editing data by temporally warping pixel pairs using optical flow fields derived from image editing samples, treating images as a special case of videos. To align the generative and editing capabilities between images and videos, they introduce a modality imitation loss, complemented by referring expression segmentation and region-aware losses operating in both latent and attention spaces. This enables the model to internalize language-guided editing without external masks or auxiliary models. Training solely on the generated data yields a unified model that significantly improves diversity and accuracy across multi-level video editing tasks.
πŸ“ Abstract
In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited task scalability. As a result, the diversity of editing tasks remains substantially narrower than that available for image editing models. We develop a pixel-pair temporal warped flow field that can directly generate corresponding video editing samples in real time from image editing samples, and we demonstrate across multiple levels of video editing tasks that a model can learn video editing using only such data. We regard the image modality as a particular form of the video modality. Accordingly, we design a modality mimic generation loss and a modality mimic editing loss to relatively align the capabilities--and thereby the output distributions--of the two modalities through mutual imitation. Moreover, language-based visual editing entails the comprehension of the editing instruction and the reference visual content, the localization of the region corresponding to that instruction within the reference visual contents, and the modification of that region alone. Existing approaches predominantly rely on external aids, such as fine-tuning an additional MLLM or explicitly supplying a mask sequence as auxiliary input during inference. In contrast, we aspire for the model to internalize this capability. To that end, we introduce sense-related tasks--for instance, referring expression segmentation--along with corresponding editing-region-aware latent-level loss and attention-level loss.
Problem

Research questions and friction points this paper is trying to address.

video editing data generation
modality mimicry
mask-free editing
language-guided visual editing
task scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

mask-free editing
pixel-pair warped flow
modality mimicry
video editing data generation
region-aware latent loss
πŸ”Ž Similar Papers
No similar papers found.