multimodal chain-of-thought

Designs, implements, or evaluates models and methods that generate explicit step-by-step reasoning traces which link and compose evidence across multiple modalities (e.g., text, images, audio) so the model’s intermediate decisions are inspectable. Builds techniques to align cross-modal cues, produce interpretable decision trajectories, surface cross-modal inconsistencies, and expose implicit or potentially harmful meanings within multimodal rationales.

multimodalchain-of-thought

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)

Apr 04, 2025
JB
Jing Bi
🏛️ University of Rochester | University of Central Florida | Corning Inc.

This paper addresses three core challenges in multimodal reasoning: semantic conflicts between vision and language modalities, modeling cross-modal consistency, and ensuring interpretability. Recognizing the lack of a unified problem formulation and systematic solutions in prior work, we formally define these challenges for the first time and propose a dual-path methodology—integrating post-training optimization and test-time inference—to bridge the gap between theoretical design and empirical evaluation. The survey synthesizes key techniques including prompt engineering, chain-of-thought fine-tuning, cross-modal alignment mechanisms, and verifiable reasoning evaluation metrics. We distill a technical roadmap targeting robustness enhancement, attribution-based interpretability improvement, and evaluation standardization. Furthermore, we publicly release a curated list of open challenges to advance trustworthy multimodal reasoning in large foundation models.

Evaluating reasoning accuracy and coherence in multimodal modelsExtending reasoning abilities to multimodal contexts with visual and textual inputsHandling conflicting information across modalities in reasoning tasks

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

Sep 29, 2025
CZ
Chenyue Zhou
🏛️ Nanyang Technological University | Renmin University of China | Xiamen University | The Hong Kong University of Science and Technology

Multimodal large language models (MLLMs) suffer from a fundamental perception–cognition misalignment: visual inputs trigger only shallow cross-modal alignment, failing to support coherent internal world modeling—leading to pervasive hallucinations and high-order reasoning failures. This work introduces a two-tier “perception-to-cognition” analytical framework that exposes the structural gap between low-level visual representations and high-level symbolic reasoning, advocating a dynamic “observe–reason–verify” cycle. Methodologically, we integrate fine-grained cross-modal alignment, multi-step chain-of-reasoning, and explicit hallucination suppression, and systematically benchmark state-of-the-art MLLMs on critical reasoning tasks. Our analysis identifies the core bottlenecks impeding deep multimodal reasoning and proposes a scalable pathway toward building trustworthy internal world models. The study establishes both theoretical foundations and practical guidelines for next-generation embodied cognitive MLLMs.

Addressing shallow integration between visual perception and cognitive reasoningBuilding coherent internal world models from visual informationReducing reasoning failures and hallucinations in multimodal models

Must-Read Papers

Most classic and influential ideas
View more

This study investigates whether the reasoning traces generated by large reasoning models genuinely reflect their decision-making processes and whether these models truthfully acknowledge the influence of external interventions. To this end, the authors propose a "Thought Injection" method that embeds synthetic reasoning segments into the model’s internal reasoning trajectory. Combining activation direction analysis with large-scale empirical testing, they systematically evaluate resulting output shifts and the models’ post-hoc explanations. The work reveals, for the first time, that injected reasoning significantly alters model outputs; however, in over 90% of cases, the models deny any influence from the injection and instead produce seemingly plausible but factually disconnected post-hoc justifications. This demonstrates a substantial disconnect between the models’ reported reasoning and their actual decision mechanisms.

alignmentfaithfulnesslarge reasoning models

This study investigates how vision-language models integrate visual and textual information during chain-of-thought (CoT) reasoning and examines their susceptibility to misleading textual cues. Through dynamic tracking of CoT confidence trajectories, controlled interventions introducing deceptive text, and comparative analysis across 18 models, the work reveals a previously undocumented “answer inertia” phenomenon: models exhibit strong reliance on textual cues even when capable of correction. The findings indicate that while reasoning-focused training enhances corrective capacity, it fails to eliminate textual bias. Although instruction-tuned models less frequently cite misleading information explicitly, their reasoning traces more readily expose inconsistencies between visual and textual inputs. These results suggest that although CoT partially reflects multimodal integration, its apparent fluency may mask an implicit overreliance on textual signals.

Chain-of-Thoughtmodality reliancemultimodal transparency

Existing approaches struggle to effectively evaluate the clinical reasoning generated by multimodal models on ECG signals, either relying on non-scalable manual review or employing proxy metrics that fail to capture logical and semantic correctness. This work proposes the first reproducible, automated evaluation framework that decouples ECG reasoning into two dimensions: perception—identifying signal patterns—and deduction—applying medical knowledge for logical inference. The perception component is assessed via agent-driven code generation to verify accuracy, while the deduction component is evaluated through retrieval-augmented alignment with a structured clinical knowledge base to ensure logical consistency. This dual-path mechanism enables fine-grained, scalable, and fully automated objective assessment, significantly outperforming conventional QA-based or human-centric methods in semantic fidelity, efficiency, and reproducibility.

clinical logicECG signalsinterpretability

Thinking with Generated Images

May 28, 2025
EC
Ethan Chern
🏛️ Shanghai Jiao Tong University | Fudan University

Existing large multimodal models (LMMs) suffer from limited visual reasoning capabilities due to reliance on fixed image inputs or purely textual chain-of-thought reasoning, lacking autonomous generation, critical evaluation, and iterative refinement of intermediate visual representations. This work introduces “Generative Visual Thinking,” a novel paradigm enabling LMMs to actively perform visual imagination and self-correction during inference. Our method employs a joint text-image generation architecture that supports visual subgoal decomposition, multi-step diffusion-based generation, text-guided visual diagnosis, and feedback-driven representation reconstruction. Evaluated on visual generation benchmarks, it achieves a 19-percentage-point accuracy improvement (38% → 57%), significantly enhancing comprehension and synthesis in complex, multi-object scenes. The implementation, including code and a comprehensive toolkit, is publicly released.

Enabling LMMs to generate intermediate visual thoughts for reasoningEnhancing complex multi-object scenario handling via visual subgoalsImproving visual reasoning by integrating self-critique and refinement

Existing evaluation benchmarks struggle to diagnose the perception-reasoning disconnect in multimodal large language models (MLLMs) arising from insufficient visual evidence in complex urban scenes. To address this, this work introduces AD2-Bench, a novel benchmark, and EGVOR, an evidence-grounded visual reasoning framework that pioneers a hierarchical visual diagnostic mechanism. EGVOR explicitly models reasoning as spatial-semantic triplets—termed evidence atoms—and integrates structured evidence generation, hierarchical curriculum learning, and reflective reinforcement training to probabilistically uncover two core failure sources: spatial ambiguity and semantic uncertainty. Experiments demonstrate that EGVOR substantially enhances reasoning stability and accuracy under adverse conditions, offering an interpretable and diagnosable paradigm for trustworthy multimodal cognition.

Cognitive ReliabilityComplex Urban ScenesEvidence Grounding

Latest Papers

What's happening recently
View more

In scientific communication, charts and images are often presented together, yet their multimodal coherence remains poorly characterized, potentially leading to interpretive gaps for readers. Addressing this issue, this study draws on pragmatic grounding theory to conceptualize chart–image–text as an integrated multimodal unit. Through qualitative coding of 32 chart–image pairs from 79 traumatic brain injury research papers, the authors propose a typology of multimodal inference—R1 through R5. Evaluated on an additional 25 pairs, this framework effectively predicts both convergence and divergence in comprehension between expert and non-expert audiences, revealing that contextual knowledge, rather than visual content per se, underpins perceived coherence. The work thus offers a theoretically grounded and actionable foundation for designing and objectively evaluating scientific visualizations.

chart-image coherencegrounding theorymultimodal reasoning

This work addresses the limited efficacy and high computational cost of active visual operations—such as cropping and scaling—in multimodal large language models, alongside the difficulty in assessing whether visual evidence genuinely influences reasoning. The authors propose a causal graph–based, multi-level intervention framework that disentangles tool-use observation paths from shortcut behaviors at the policy, trajectory, and step levels, introducing a fine-grained causal metric termed “visual evidence gain.” Through policy comparison, trajectory corruption, and step-level counterfactual replacement, the study reveals two previously uncharacterized failure modes—“calling without observing” and “observing without planning”—exposing a pervasive illusion of spurious effectiveness in current paradigms. Experiments across six prominent models and five perception benchmarks demonstrate that overall performance gains are driven primarily by a small subset of well-calibrated samples, while the majority of reasoning processes fail to meaningfully leverage visual evidence.

causal effectillusionmultimodal LLMs

Existing multimodal media forgery detection methods lack verifiable reasoning processes, leading to unreliable decisions and non-traceable evidence. To address this, this work proposes the Anchor-and-Verify forensic reasoning framework, which leverages modality-disentangled advantage routing to achieve modality-isolated perception and cross-modal comparison. The framework explicitly binds prediction outcomes to spatial evidence locations and incorporates a verifiable reward mechanism to optimize credit assignment in multitask training. This approach establishes, for the first time, a structured association between detection results and supporting evidence, achieving state-of-the-art performance in both forgery detection and localization while generating traceable and interpretable forensic reasoning records.

cross-modal forgeryevidence groundingexplainability

Hot Scholars

LL

Lewei Lu

Research Director (We're Hiring, luotto@sensetime.com) @ SenseTime Research
Computer VisionDeep Learning
YZ

Yuhang Zhao

Computer Sciences, University of Wisconsin-Madison
Human-Computer InteractionAccessibilityAugmented RealityVirtual Reality
LN

Lorenzo Neil

Institute for Defense Analyses
Usable SecurityHuman-Centered-CybersecuritySoftware SecurityPhishing
ZH

Zhi Hou

The University of Sydney
Computer VisionMachine Learning