🤖 AI Summary
Existing drag-based editing methods struggle to balance manipulation precision with generation naturalness. This work proposes MoRe-Drag, which introduces a novel motion-as-evidence latent recomposition mechanism. Specifically, it injects pixel-space warping as a motion prior into the diffusion model sampling trajectory and integrates region-aware latent recomposition with stage-adaptive conditioning to achieve precise editing. Furthermore, the framework incorporates multimodal large language model adaptation to support instruction-free interaction. Extensive evaluations on the DragBench benchmark demonstrate that MoRe-Drag significantly outperforms existing state-of-the-art methods, effectively enhancing dragging accuracy while preserving semantic consistency and visual realism.
📝 Abstract
Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, inpainting, and anchor regions, coupled with stage-adaptive conditioning that progressively shifts from motion-grounded structure formation to semantic refinement. We further support an instruction-free interface by adapting the MLLM-based text encoder for drag-aware instruction inference. Experiments on DragBench-SR and DragBench-DR show that MoRe-Drag substantially improves drag precision over strong base editors and achieves superior drag accuracy among SOTA drag-based methods, while delivering strong semantic consistency and visually realistic results. Code and dataset will be publicly released.