Score
Designs and evaluates chain-of-thought reasoning processes and datasets that explicitly represent and supervise selection and inspection of spatial regions in visual inputs. This involves creating annotation schemes, training signals, and coarse-to-fine inference steps that teach when and where to crop or focus and how to use local evidence from image regions in multi-step reasoning.
Current multimodal chain-of-thought (MCoT) reasoning lacks a systematic survey, suffering from conceptual ambiguity, methodological fragmentation, and incomplete modality coverage. This work establishes the first unified taxonomy and comprehensive classification framework for MCoT, encompassing six modalities: image, video, speech, 3D, structured data, and cross-modal combinations. We rigorously formalize foundational definitions and synthesize core methodologies—including cross-modal alignment, stepwise multimodal reasoning modeling, interpretability analysis, and task-driven evaluation. Furthermore, we identify critical pathways toward multimodal artificial general intelligence (AGI) and articulate key open challenges. The survey critically analyzes over 200 seminal works, offering a methodological guide and technology roadmap for high-impact domains such as robotics, healthcare, and autonomous driving. To our knowledge, this is the first authoritative, comprehensive survey dedicated to MCoT.
Existing text-based chain-of-thought (CoT) reasoning treats visual input as static context, creating a “semantic gap” between perception and symbolic reasoning. To bridge this gap, we propose a paradigm shift—from “thinking about images” to “thinking with images”—and introduce the first systematic three-stage framework wherein vision serves as a dynamic cognitive workspace: visual input → manipulable intermediate representation → autonomous generation and operation. Our method integrates programmable visual operations, intrinsic imagination mechanisms, cross-modal alignment, and textual CoT, elevating vision from passive input to an active medium for reasoning. This framework establishes the theoretical foundation for “thinking with images,” advances multimodal AI toward human-like cognitive autonomy, and provides a clear roadmap for evaluation design, key technology validation, and future research directions.
This study addresses a critical yet overlooked issue in multimodal large language models: their performance significantly degrades on visual spatial reasoning tasks when chain-of-thought (CoT) prompting is applied. Through a systematic evaluation of 17 models across 13 spatial reasoning benchmarks, the work reveals for the first time that CoT induces a detrimental effect—models overly rely on textual priors while neglecting visual input, leading to increased hallucination and reduced accuracy. The authors introduce the No-Image++ ablation method to demonstrate that models exhibit severe shortcut learning, prioritizing spurious textual cues over genuine visual evidence. These findings not only expose the fragility of current multimodal reasoning systems but also provide a crucial diagnostic tool to guide the development of more robust and grounded multimodal architectures.
Existing visual reasoning research is fragmented across subdomains—including relational, symbolic, temporal, causal, and commonsense reasoning—lacking a unified taxonomy, comparable evaluation protocols, and systematic analysis. Method: We propose the first cross-paradigmatic unified classification framework for visual reasoning, integrating graph neural networks, memory-augmented architectures, attention mechanisms, and neuro-symbolic methods into a cohesive perception-reasoning architecture. We further design a multidimensional evaluation protocol assessing functional correctness, structural consistency, and causal validity. Contribution/Results: Our analysis reveals shared bottlenecks across state-of-the-art methods—particularly in out-of-distribution generalization, interpretability, and performance under weak supervision. The framework establishes a theoretical benchmark and practical technical guidelines for visual reasoning, enabling robust, trustworthy AI applications in domains such as autonomous driving and medical diagnosis.
This study investigates how chain-of-thought (CoT) design influences the generalization capability of vision-language models (VLMs) on vision-centric reasoning tasks. Using a controlled maze-solving benchmark, we systematically compare three CoT paradigms—language-descriptive, coordinate-based, and visual-operational—implemented on Qwen2.5-VL-7B and trained via supervised fine-tuning followed by reinforcement learning to auto-generate intermediate reasoning steps. Results reveal that the minimal coordinate-based CoT—retaining only essential spatial position information—achieves superior cross-scale generalization, exhibiting a “shorter is stronger” effect: it converges faster and attains significantly higher final accuracy than verbose language-based or visual-operational CoTs. This challenges the prevailing assumption that longer reasoning chains inherently yield better performance, and provides the first empirical evidence that concise, spatially grounded representations constitute a fundamental design principle for enhancing VLM generalization in visual reasoning.
This work addresses two key challenges in multi-image fine-grained visual reasoning: reliance on human-annotated question-answer pairs and difficulty modeling cross-image logical relationships. Methodologically, we propose a self-supervised chain-of-reasoning framework that constructs image triplets to uncover intrinsic visual constraints, employs chain-of-thought prompting for stepwise inference, and integrates rule-guided reinforcement learning to encourage the model to autonomously attend to subtle visual differences and perform interpretable logical deduction—entirely without annotated QA pairs. Our primary contribution is the first integration of self-supervised contrastive learning with structured chain-of-reasoning, enabling fine-grained cross-image comparison and generalization to complex logical patterns. Experiments demonstrate significant improvements over existing unsupervised methods on multi-image reasoning benchmarks, while also exhibiting strong transferability to general vision tasks.
Existing vision-language models (VLMs) lack self-reflection and error-correction capabilities in multi-turn visual reasoning. Method: We propose a verifiable, self-reflective multi-step visual reasoning framework that actively invokes external tools for image analysis—rather than passively perceiving inputs—and employs a redundancy-penalized reinforcement learning (RL) strategy to encourage multi-scale exploration and trajectory-level self-correction. We further construct a challenging, answer-verifiable multi-turn visual question-answering dataset. Our approach integrates high-resolution image inputs, cold-start supervised fine-tuning (SFT), and redundancy-aware RL to support iterative tool invocation and reasoning trajectory assessment. Contribution/Results: Experiments demonstrate significant improvements in multi-step reasoning accuracy, robustness, and self-correction capability across multiple visual understanding benchmarks, establishing a new paradigm for trustworthy visual reasoning.
This work addresses the limitations of existing vision-language models in robust spatial reasoning, particularly their inadequate credit assignment and lack of depth perception. To overcome these challenges, the authors propose SCOUT, a novel framework that integrates structured chain-of-thought (CoT) reasoning with explicit 3D environmental modeling and a reinforcement learning algorithm featuring multi-target process rewards to enable fine-grained credit assignment, complemented by a tailored advantage estimation method. Additionally, they introduce SCOUT-24k, the first structured dataset for spatial reasoning. Experimental results demonstrate that SCOUT-3B achieves performance gains of 16.85% and 6.3% on general and complex spatial tasks, respectively, while SCOUT-7B surpasses GPT-4o by 4.28% and exhibits strong generalization across multi-image and video scenarios.
This work addresses the limitations of existing cross-modal chain-of-thought methods, which often over-rely on a single coarse-grained image region and suffer from semantic discontinuities across reasoning steps in complex visual reasoning tasks. To overcome these issues, we propose CoCoT, a novel framework that introduces a dynamic multi-region visual focusing mechanism coupled with relation-aware reasoning to enable collaborative multi-region integration and construct coherent cross-modal reasoning chains. Additionally, we curate CoCoT-70K, a high-quality dataset comprising 74,691 samples. Extensive experiments demonstrate that CoCoT achieves substantial performance gains across six challenging benchmarks, improving average accuracy by 15.4% on LLaVA-1.5 and by 4.0% on Qwen2-VL.
This study investigates the impact of external information—such as spatial cues, commonsense knowledge, and chain-of-thought prompts—on visual spatial reasoning (VSR) performance. Through hypothesis-driven controlled experiments on two public benchmarks, the authors systematically evaluate three categories of vision-language models. Their findings reveal that a single, precise spatial cue consistently outperforms multi-context fusion; weakly relevant or excessive commonsense knowledge degrades performance; and chain-of-thought prompting is beneficial only when spatial localization is sufficiently accurate. This work is the first to demonstrate that “more information is not always better” in VSR, advocating instead for the selective injection of task-aligned signals and clarifying the conditions under which spatial localization and chain-of-thought reasoning synergistically enhance performance.
This work addresses the lack of explicit grounding between reasoning steps and visual evidence in existing vision-language models, which hinders verifiability and supervision. To overcome this limitation, the authors propose a visually anchored reasoning mechanism that explicitly binds each natural language inference step to specific point- or box-level regions in the image, enabling transparent and supervisable multimodal reasoning. The approach integrates synthetic data distillation, SAM3-based automatic annotation, and a reinforcement learning strategy that jointly optimizes answer correctness and visual alignment. Built upon the Gemma3 model family and fine-tuned accordingly, the 4B-parameter variant significantly outperforms non-anchored baselines on two counting and four spatial reasoning tasks, achieving performance comparable to or even exceeding that of the much larger 27B-parameter counterpart within the same model series.