region-aware chain-of-thought

Designs and evaluates chain-of-thought reasoning processes and datasets that explicitly represent and supervise selection and inspection of spatial regions in visual inputs. This involves creating annotation schemes, training signals, and coarse-to-fine inference steps that teach when and where to crop or focus and how to use local evidence from image regions in multi-step reasoning.

region-awarechain-of-thought

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Jun 30, 2025
ZS
Zhaochen Su
🏛️ The Hong Kong University of Science and Technology | UNC-Chapel Hill | Microsoft | The Chinese University of Hong Kong | UIUC

Existing text-based chain-of-thought (CoT) reasoning treats visual input as static context, creating a “semantic gap” between perception and symbolic reasoning. To bridge this gap, we propose a paradigm shift—from “thinking about images” to “thinking with images”—and introduce the first systematic three-stage framework wherein vision serves as a dynamic cognitive workspace: visual input → manipulable intermediate representation → autonomous generation and operation. Our method integrates programmable visual operations, intrinsic imagination mechanisms, cross-modal alignment, and textual CoT, elevating vision from passive input to an active medium for reasoning. This framework establishes the theoretical foundation for “thinking with images,” advances multimodal AI toward human-like cognitive autonomy, and provides a clear roadmap for evaluation design, key technology validation, and future research directions.

Bridging semantic gap between vision and symbolic reasoningDeveloping dynamic visual cognitive workspace in modelsEvolving AI from thinking about to thinking with images

This study addresses a critical yet overlooked issue in multimodal large language models: their performance significantly degrades on visual spatial reasoning tasks when chain-of-thought (CoT) prompting is applied. Through a systematic evaluation of 17 models across 13 spatial reasoning benchmarks, the work reveals for the first time that CoT induces a detrimental effect—models overly rely on textual priors while neglecting visual input, leading to increased hallucination and reduced accuracy. The authors introduce the No-Image++ ablation method to demonstrate that models exhibit severe shortcut learning, prioritizing spurious textual cues over genuine visual evidence. These findings not only expose the fragility of current multimodal reasoning systems but also provide a crucial diagnostic tool to guide the development of more robust and grounded multimodal architectures.

Chain-of-ThoughtHallucinationMultimodal LLMs

Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies

Aug 14, 2025
AS
Ayushman Sarkar
🏛️ Birbhum Institute of Engineering and Technology | Universiti Malaya

Existing visual reasoning research is fragmented across subdomains—including relational, symbolic, temporal, causal, and commonsense reasoning—lacking a unified taxonomy, comparable evaluation protocols, and systematic analysis. Method: We propose the first cross-paradigmatic unified classification framework for visual reasoning, integrating graph neural networks, memory-augmented architectures, attention mechanisms, and neuro-symbolic methods into a cohesive perception-reasoning architecture. We further design a multidimensional evaluation protocol assessing functional correctness, structural consistency, and causal validity. Contribution/Results: Our analysis reveals shared bottlenecks across state-of-the-art methods—particularly in out-of-distribution generalization, interpretability, and performance under weak supervision. The framework establishes a theoretical benchmark and practical technical guidelines for visual reasoning, enabling robust, trustworthy AI applications in domains such as autonomous driving and medical diagnosis.

Challenges in scalability, integration, benchmarks, and weak supervision in reasoningLack of unified analysis across visual reasoning types and methodologiesNeed for systematic evaluation of reasoning models' correctness and limitations

Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization

Nov 27, 2025
YD
Yifan Du
🏛️ Renmin University of China | ByteDance Seed

This study investigates how chain-of-thought (CoT) design influences the generalization capability of vision-language models (VLMs) on vision-centric reasoning tasks. Using a controlled maze-solving benchmark, we systematically compare three CoT paradigms—language-descriptive, coordinate-based, and visual-operational—implemented on Qwen2.5-VL-7B and trained via supervised fine-tuning followed by reinforcement learning to auto-generate intermediate reasoning steps. Results reveal that the minimal coordinate-based CoT—retaining only essential spatial position information—achieves superior cross-scale generalization, exhibiting a “shorter is stronger” effect: it converges faster and attains significantly higher final accuracy than verbose language-based or visual-operational CoTs. This challenges the prevailing assumption that longer reasoning chains inherently yield better performance, and provides the first empirical evidence that concise, spatially grounded representations constitute a fundamental design principle for enhancing VLM generalization in visual reasoning.

Compares language, grounding, and visual CoT formats in maze-solvingEvaluates Chain-of-Thought designs for visual reasoning generalizationFinds concise, essential grounding CoT generalizes best across tasks

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

Jun 27, 2025
XC
Xi Chen
🏛️ HKU | CUHK | Tongyi Lab | Alibaba Group | HUST

This work addresses two key challenges in multi-image fine-grained visual reasoning: reliance on human-annotated question-answer pairs and difficulty modeling cross-image logical relationships. Methodologically, we propose a self-supervised chain-of-reasoning framework that constructs image triplets to uncover intrinsic visual constraints, employs chain-of-thought prompting for stepwise inference, and integrates rule-guided reinforcement learning to encourage the model to autonomously attend to subtle visual differences and perform interpretable logical deduction—entirely without annotated QA pairs. Our primary contribution is the first integration of self-supervised contrastive learning with structured chain-of-reasoning, enabling fine-grained cross-image comparison and generalization to complex logical patterns. Experiments demonstrate significant improvements over existing unsupervised methods on multi-image reasoning benchmarks, while also exhibiting strong transferability to general vision tasks.

Enable CoT reasoning across multiple visual cuesLearn reasoning via self-supervised image comparisonsOvercome reliance on manual QA pairs for VLMs

Latest Papers

What's happening recently
View more

Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images

Dec 19, 2025
WY
Wenhao Yang
🏛️ Nanjing University | Alibaba Group | Zhejiang University

Existing vision-language models (VLMs) lack self-reflection and error-correction capabilities in multi-turn visual reasoning. Method: We propose a verifiable, self-reflective multi-step visual reasoning framework that actively invokes external tools for image analysis—rather than passively perceiving inputs—and employs a redundancy-penalized reinforcement learning (RL) strategy to encourage multi-scale exploration and trajectory-level self-correction. We further construct a challenging, answer-verifiable multi-turn visual question-answering dataset. Our approach integrates high-resolution image inputs, cold-start supervised fine-tuning (SFT), and redundancy-aware RL to support iterative tool invocation and reasoning trajectory assessment. Contribution/Results: Experiments demonstrate significant improvements in multi-step reasoning accuracy, robustness, and self-correction capability across multiple visual understanding benchmarks, establishing a new paradigm for trustworthy visual reasoning.

Addresses self-reflection and correction in visual reasoning trajectoriesEnables deep, reliable multi-turn reasoning with imagesImproves performance on complex visual understanding benchmarks

This work addresses the limitations of existing vision-language models in robust spatial reasoning, particularly their inadequate credit assignment and lack of depth perception. To overcome these challenges, the authors propose SCOUT, a novel framework that integrates structured chain-of-thought (CoT) reasoning with explicit 3D environmental modeling and a reinforcement learning algorithm featuring multi-target process rewards to enable fine-grained credit assignment, complemented by a tailored advantage estimation method. Additionally, they introduce SCOUT-24k, the first structured dataset for spatial reasoning. Experimental results demonstrate that SCOUT-3B achieves performance gains of 16.85% and 6.3% on general and complex spatial tasks, respectively, while SCOUT-7B surpasses GPT-4o by 4.28% and exhibits strong generalization across multi-image and video scenarios.

3D understandingcredit assignmentdepth perception

This work addresses the limitations of existing cross-modal chain-of-thought methods, which often over-rely on a single coarse-grained image region and suffer from semantic discontinuities across reasoning steps in complex visual reasoning tasks. To overcome these issues, we propose CoCoT, a novel framework that introduces a dynamic multi-region visual focusing mechanism coupled with relation-aware reasoning to enable collaborative multi-region integration and construct coherent cross-modal reasoning chains. Additionally, we curate CoCoT-70K, a high-quality dataset comprising 74,691 samples. Extensive experiments demonstrate that CoCoT achieves substantial performance gains across six challenging benchmarks, improving average accuracy by 15.4% on LLaVA-1.5 and by 4.0% on Qwen2-VL.

chain-of-thoughtcross-modal reasoningmulti-region grounding

This study investigates the impact of external information—such as spatial cues, commonsense knowledge, and chain-of-thought prompts—on visual spatial reasoning (VSR) performance. Through hypothesis-driven controlled experiments on two public benchmarks, the authors systematically evaluate three categories of vision-language models. Their findings reveal that a single, precise spatial cue consistently outperforms multi-context fusion; weakly relevant or excessive commonsense knowledge degrades performance; and chain-of-thought prompting is beneficial only when spatial localization is sufficiently accurate. This work is the first to demonstrate that “more information is not always better” in VSR, advocating instead for the selective injection of task-aligned signals and clarifying the conditions under which spatial localization and chain-of-thought reasoning synergistically enhance performance.

commonsense knowledgeinformation injectionspatial grounding

This work addresses the lack of explicit grounding between reasoning steps and visual evidence in existing vision-language models, which hinders verifiability and supervision. To overcome this limitation, the authors propose a visually anchored reasoning mechanism that explicitly binds each natural language inference step to specific point- or box-level regions in the image, enabling transparent and supervisable multimodal reasoning. The approach integrates synthetic data distillation, SAM3-based automatic annotation, and a reinforcement learning strategy that jointly optimizes answer correctness and visual alignment. Built upon the Gemma3 model family and fine-tuned accordingly, the 4B-parameter variant significantly outperforms non-anchored baselines on two counting and four spatial reasoning tasks, achieving performance comparable to or even exceeding that of the much larger 27B-parameter counterpart within the same model series.

countingreasoning tracesspatial reasoning

Hot Scholars

XH

Xiaoguang Han

Assistant Professor, The Chinese University of Hong Kong, Shenzhen
Computer VisionComputer Graphics
BL

Bang Liu

Associate Professor at the University of Montreal, Canada CIFAR AI Chair at Mila
Natural Language ProcessingDeep LearningMachine LearningData Mining
JZ

Jinhua Zhao

Professor of Cities and Transportation, Massachusetts Institute of Technology
Urban MobilityTravel BehaviorTransportation PolicyPublic Transit
YF

Yixu Feng

Northwestern Polytechnical University
Artificial IntelligenceComputer VisionLow-level Vision
HH

Heye Huang

University of Wisconsin–Madison
Autonomous SystemsMulti-AgentsRisk AssessmentInteractive Decision-Making