Beyond Plausibility: Verifiable Fine-Grained Image Editing on Structured Assets

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the evaluation gap between visual realism and verifiability in fine-grained image editing, as well as the subjective uncertainty inherent in human or VLM-based assessments. To this end, it introduces VeriEdit-Bench, the first deterministic evaluation benchmark grounded in structured asset source code. Methodologically, SVG source code is leveraged to control the editing process, generating precise targets and pixel-level masks that enable quantitative evaluation across four axes, including fidelity and preservation rate. This benchmark eliminates subjective bias by providing a fully reproducible scoring mechanism. Furthermore, it reveals significant performance disparities among existing models across multiple dimensions, cross-scenario ranking inversions, and distinct capability failure modes.
📝 Abstract
Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically rely on human or vision--language model judgments, while deterministic evaluation has largely focused on synthetic shape canvases, with application-oriented extensions primarily limited to charts. This makes it difficult to determine precisely how much of a requested edit was executed, where unintended changes occurred, and whether small differences between models reflect genuine editing capability or evaluator uncertainty. To bridge this gap, we present VeriEdit-Bench, a benchmark for fine-grained, instruction-faithful image editing across realistic structured assets with deterministic, four-axis evaluation. Its 1,740 cases are compiled from the source code of 153 Scalable Vector Graphics (SVG) graphics, charts, web interfaces, and presentation slides. Controlled source-code edits preserve the original visual context while yielding exact target images, pixel-level edit masks, and explicit edit specifications, enabling reproducible scoring along four axes: edit fidelity, preservation, localization, and magnitude. Evaluating eleven editors, we find that even the strongest model remains far from full credit; rankings for the same recoloring operation reverse between charts and SVG graphics; and outputs with similar pixel-accuracy profiles can still differ substantially in localization and change magnitude. This decomposition yields graded, verifiable feedback and exposes model-specific capability and failure profiles that holistic scores or evaluator-dependent judgments may obscure.
Problem

Research questions and friction points this paper is trying to address.

fine-grained image editing
verifiable evaluation
benchmark
structured assets
edit fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

fine-grained image editing
verifiable benchmark
structured assets
deterministic evaluation
SVG source code
💼 Related Jobs
No related jobs found.