Scalable Minimal-Change Learning for Controllable Image Editing

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of unintended modifications in existing instruction-based image editing, which hinders minimal-change objectives. To this end, we propose a reinforcement learning-based optimization paradigm for minimal editing. Specifically, an agentic vision-language model is employed to audit both editing requests and unintended alterations, while a group-level scoring mechanism—requiring no human annotation—is designed to provide consistent reward signals for FLUX.1 Kontext-dev. Experimental results demonstrate that the proposed method improves EditScore to 5.88 and reduces non-target pixel changes by 8.4%. Furthermore, blind user evaluations confirm its effectiveness, achieving high-fidelity and precise image editing with minimal unintended modifications.
📝 Abstract
Image editing should change only the attributes specified by an instruction while preserving everything else, yet current methods often make unintended changes. We treat this minimal-change principle as an optimization objective for instruction-based editing. Latent L1 regularization is a poor proxy for output locality in modern nonlinear generators and often requires supervision unavailable at scale. We instead optimize edit outcomes with reinforcement learning. An agentic vision-language reward model audits each source image, instruction, and edited image for two failure types: unimplemented requested changes and unintended changes. A group-level rubric merges and verifies these issues to provide consistent rewards across candidate edits without per-instruction human annotations. On FLUX.1 Kontext-dev, ARRO raises average EditScore from 5.21 to 5.88 across MinEval, MagicBrush, AnyBench, and Emu-Edit. On 600 evaluation examples, it reduces off-target pixel change by 8.4% relative to the base editor. Reward and SFT controls, blinded human evaluations, and transfer to OmniGen2 provide complementary evidence. Code: https://github.com/Showwwwwwwww/ARRO
Problem

Research questions and friction points this paper is trying to address.

controllable image editing
minimal-change principle
instruction-based editing
unintended changes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Minimal-Change Editing
Vision-Language Reward Model
Controllable Image Editing
Scalable Optimization
🔎 Similar Papers
No similar papers found.