Rubric-CEPR: Self-Evolving Image Editing via Reward-Verified Self-Distillation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of existing image editing methods on costly manual annotations or external reward models, which often erroneously reward superficially plausible yet substantively failed samples. To this end, this work proposes a self-evolutionary framework that introduces a Contrastive Editing Preservation Reward (CEPR) based on internal representations, combined with a non-compensatory gating mechanism for candidate filtering. The framework achieves iterative optimization without external supervision through planner-based instruction generation, multi-candidate sampling, frozen critic scoring, and lightweight adapter self-distillation. Experimental results demonstrate that the proposed method improves the ImgEdit metric by 5.5% and object isolation performance by 24.9% on Qwen-Image-Edit, while yielding a 7.8% improvement on Step1X-Edit.
📝 Abstract
Instruction-guided image editors have become highly capable, yet improving them further still depends on human-edited training pairs or external reward models. Such supervision is costly to obtain and can reward plausible failures: a realistic output may leave the requested change undone or alter content that should be preserved. In this work, we strive to improve a pretrained image editor using only its own generations, without human-edited targets or an external training-time reward model. To this end, we propose a self-evolving framework, named Rubric-CEPR, that verifies the editor's own samples with its internal representations through a rubric-augmented Contrastive Edit-Preservation Reward (CEPR). A Planner proposes structured edit instructions from unlabeled images, the Editor samples multiple candidate edits, and a frozen Critic scores each candidate with decomposed rubric checks for edit realization, removal of the old state, and content preservation, using features already exposed by the editor. Non-compensatory gates reject infeasible candidates, and the best verified candidate is distilled into the editor through lightweight adapter training. On Qwen-Image-Edit, Rubric-CEPR improves ImgEdit from 4.36 to 4.60 (+5.5%), with a +24.9% gain on object isolation, and transfers to GEdit-Bench and Complex-Edit. The same procedure also improves Step1X-Edit by +7.8% on ImgEdit. We hope our approach will serve as a solid baseline for image editors that improve themselves from their own verified samples. Our code is publicly available at $\href{https://riteshthawkar.github.io/Rubric-CEPR/}{\text{this URL}}$
Problem

Research questions and friction points this paper is trying to address.

image editing
self-evolution
self-distillation
reward verification
instruction-guided editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Image Editing
Reward Verification
Contrastive Edit-Preservation Reward
Self-Evolving Framework
🔎 Similar Papers
No similar papers found.