🤖 AI Summary
This study addresses the limitation of existing vision-language models (VLMs) in incorporating images into reasoning chains, as well as the lack of interpretability in omni-modal models relying on rasterized representations. To this end, we propose SVGLM, a framework that leverages Scalable Vector Graphics (SVG) to bridge text and imagery by exploiting their dual nature as both image descriptions and textual instructions. By constructing a large-scale SVG editing dataset and fine-tuning open-source VLMs, our framework empowers general-purpose models to generate images during reasoning, achieving a compact and interpretable unified paradigm for visual-textual reasoning. Experiments demonstrate that SVGLM exhibits strong SVG generation capabilities and "thinking with images" intelligence on mathematical reasoning benchmarks, effectively bridging the gap between textual reasoning and pixel-level imagery.
📝 Abstract
Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorporating images into the reasoning chain. Though Omnimodal models have made efforts in unifying text and image generation, they focus on visual tasks in the open-domain, lacking tractability due to rasterized or latent representations of images. We introduce SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks. We exploit the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generating images within the reasoning process. We provide a large curated dataset of SVG-based image editing dataset, as well as the paradigm to tune open-source VLMs. Experiments on a mathematical reasoning benchmark demonstrate that SVGLM achieves strong SVG generation power as well as think-with-image intelligence. Our results highlight SVG as a suitable medium for building more robust digital domain agents, bridging the gap between text-based thinking and pixel-based images.