SolveEdit: Benchmarking Visual Problem Solving in Generative Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks in evaluating models' ability to preserve irrelevant content while precisely executing scene transformations during goal-driven visual problem solving. To this end, we propose the SolveEdit benchmark. First, we design an atomic transformation contract scoring system that eliminates the need for a single reference output, enabling quantitative assessment of both task completion and unintended modifications. Second, we construct a two-stage visual planning architecture that instantiates transformations prior to generation, thereby enhancing the reasoning capabilities of generative models. Experimental results demonstrate that the strongest baseline achieves only 57.0% on SolveEdit. Incorporating the proposed planner yields an average improvement of 9.1 points, with GPT-Image-2 reaching 71.6%, validating the effectiveness of our approach.
📝 Abstract
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing that change without disturbing unrelated content. However, existing benchmarks mainly evaluate perception, generation, or explicitly specified transformations, leaving goal-driven visual problem solving underexplored. To bridge this gap, we introduce SolveEpIT, a benchmark for visual problem solving through scene transformation. Given an image and a goal, a model must infer a valid transformation from the request, the scene, or a visually expressed rule, then execute it while preserving unrelated content. SoLvEEDrr contains 2,728 cases. Atomic transition contracts specify required and protected conditions, enabling SoLvEScoRE to measure completion and unintended changes without a single reference output. The strongest evaluated model achieves only57.0% SolvEScore. We further introduce SolveEdiT-PLAN, a two-stage visual planner that instantiates the transition before generation. Under matched single-generation evaluation, it improves SoLvEScoRE by 9.1 points on average across three tested generators, including a gain from 57.0% to 71.6% for GPT-Image-2, without modifying the editor.
Problem

Research questions and friction points this paper is trying to address.

Visual Problem Solving
Scene Transformation
Benchmarking
Generative Models
Goal-driven Editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Problem Solving
Benchmark
Scene Transformation
Atomic Transition Contracts
Two-stage Visual Planner