RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation

📅 2026-07-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing image-to-SVG generation methods, which rely on single-pass, open-loop inference and often suffer from geometric misalignment, error accumulation, and visual hallucinations, particularly on complex images. To overcome these issues, the authors propose RefineSVG, a novel framework that introduces the first single-step closed-loop visual feedback mechanism. It leverages an external SVG renderer to produce a Diff-Map residual image as a feedback signal and employs a ReAct-style refinement strategy to iteratively guide a multimodal large language model toward improved outputs. Additionally, the method incorporates an SVG-tailored semantic vocabulary that reduces code length by over 52%. Experiments demonstrate that RefineSVG significantly outperforms current baselines in reconstruction fidelity, structural accuracy, and code efficiency.
📝 Abstract
We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation through self-correction. Existing MLLM-based approaches rely on single-pass open-loop inference, where the model receives visual input only once and must generate thousands of SVG code tokens without intermediate verification. This paradigm inevitably leads to geometric drift, error accumulation, and visual hallucination on complex images. RefineSVG overcomes this limitation by invoking an external rendering engine after an initial SVG generation pass to compare the rendered output against the target image. The comparison yields a multi-dimensional visual residual map (Diff-Map) that is fed back to the model as a ReAct-style correction signal, driving a targeted correction step. To support this render-observe-correct interaction, we further introduce an SVG-oriented semantic vocabulary that compresses token sequences by over 52%. A progressive training pipeline spanning supervised fine-tuning, rejection-sampling cold-start data construction, and end-to-end agentic reinforcement learning aligns the model with closed-loop visual correction. Extensive experiments show that RefineSVG consistently outperforms existing baselines in reconstruction fidelity, structural accuracy, and code efficiency.Code is available at https://github.com/liuxiaobo66/RefineSVG.
Problem

Research questions and friction points this paper is trying to address.

image-to-SVG generation
geometric drift
error accumulation
visual hallucination
open-loop inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual feedback
closed-loop reinforcement learning
image-to-SVG generation
Diff-Map
SVG semantic vocabulary