DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning

📅 2025-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing vision-language models (VLMs) predominantly rely on text-dominated reasoning paradigms, limiting their capacity for fine-grained, truly image-centric multimodal reasoning—i.e., “thinking in images.” Method: We propose an image-interactive reasoning paradigm, introducing the first large-scale, interleaved image-text dataset with 31K chain-of-thought trajectories, and design an endogenous visual thinking generation mechanism. This mechanism performs reasoning directly within the visual embedding space—without external tool invocation—enabling more flexible and image-faithful thought evolution. Contribution/Results: Experiments demonstrate substantial improvements over strong baselines across multiple multimodal reasoning benchmarks. Our results validate both the efficacy of the curated dataset construction strategy and the superiority of visual-space reasoning modeling. The proposed framework advances VLMs toward deeper visual understanding by shifting reasoning from textual abstraction to intrinsic visual representation, offering a novel pathway for next-generation multimodal intelligence.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
The "thinking with images" paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of-thought to image-interactive reasoning. By invoking visual tools or generating intermediate visual representations, VLMs can iteratively attend to fine-grained regions, enabling deeper image understanding and more faithful multimodal reasoning. As an emerging paradigm, however, it still leaves substantial room for exploration in data construction accuracy, structural design, and broader application scenarios, which offer rich opportunities for advancing multimodal reasoning. To further advance this line of work, we present DeepSketcher, a comprehensive suite comprising both an image-text interleaved dataset and a self-contained model. The dataset contains 31k chain-of-thought (CoT) reasoning trajectories with diverse tool calls and resulting edited images, covering a wide range of data types and manipulation instructions with high annotation accuracy. Building on this resource, we design a model that performs interleaved image-text reasoning and natively generates "visual thoughts" by operating directly in the visual embedding space, rather than invoking external tools and repeatedly re-encoding generated images. This design enables tool-free and more flexible "thinking with images". Extensive experiments on multimodal reasoning benchmarks demonstrate strong performance, validating both the utility of the dataset and the effectiveness of the model design.
Problem

Research questions and friction points this paper is trying to address.

Advancing multimodal reasoning with visual manipulation tools
Creating accurate dataset for image-text chain-of-thought reasoning
Developing model that generates visual thoughts without external tools
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates visual thoughts in embedding space
Uses tool-free image-text interleaved reasoning
Provides dataset with diverse manipulation trajectories
C
Chi Zhang
School of Computer Science, Wuhan University
Haibo Qiu
Haibo Qiu
University of Sydney
Multimodal LLMVision and LanguageComputer Vision
Q
Qiming Zhang
The University of Sydney
Z
Zhixiong Zeng
Meituan Inc
L
Lin Ma
Meituan Inc
J
Jing Zhang
School of Computer Science, Wuhan University