Score
Design and apply methods and metrics that decompose and evaluate the stepwise reasoning chains of multimodal systems by tracing modality-specific token contributions through intermediate steps and locating where and how visual or other modality content is integrated. Build diagnostics to quantify information sufficiency and focus at each reasoning step, detect loss of attention or coherence across extra compute or iterations, and identify points where errors or missing information arise.
This paper addresses three core challenges in multimodal reasoning: semantic conflicts between vision and language modalities, modeling cross-modal consistency, and ensuring interpretability. Recognizing the lack of a unified problem formulation and systematic solutions in prior work, we formally define these challenges for the first time and propose a dual-path methodology—integrating post-training optimization and test-time inference—to bridge the gap between theoretical design and empirical evaluation. The survey synthesizes key techniques including prompt engineering, chain-of-thought fine-tuning, cross-modal alignment mechanisms, and verifiable reasoning evaluation metrics. We distill a technical roadmap targeting robustness enhancement, attribution-based interpretability improvement, and evaluation standardization. Furthermore, we publicly release a curated list of open challenges to advance trustworthy multimodal reasoning in large foundation models.
Current multimodal chain-of-thought (MCoT) reasoning lacks a systematic survey, suffering from conceptual ambiguity, methodological fragmentation, and incomplete modality coverage. This work establishes the first unified taxonomy and comprehensive classification framework for MCoT, encompassing six modalities: image, video, speech, 3D, structured data, and cross-modal combinations. We rigorously formalize foundational definitions and synthesize core methodologies—including cross-modal alignment, stepwise multimodal reasoning modeling, interpretability analysis, and task-driven evaluation. Furthermore, we identify critical pathways toward multimodal artificial general intelligence (AGI) and articulate key open challenges. The survey critically analyzes over 200 seminal works, offering a methodological guide and technology roadmap for high-impact domains such as robotics, healthcare, and autonomous driving. To our knowledge, this is the first authoritative, comprehensive survey dedicated to MCoT.
This paper addresses core challenges in multimodal AI systems under open, uncertain environments—namely, weak perception, shallow reasoning, deficient planning, and poor generalization. To tackle these, it proposes a four-stage evolutionary paradigm: “Perception → Reasoning → Reflection → Planning,” systematically mapping the technical trajectory of multimodal reasoning in the large-model era. It introduces the novel concept of Native Large Multimodal Reasoning Models (N-LMRMs), emphasizing scalability, embodiment, and autonomous planning capability, integrated with Multimodal Chain-of-Thought (MCoT), cross-modal alignment, instruction tuning, and multimodal reinforcement learning. The work identifies three fundamental bottlenecks: full-modality generalization, deep reasoning, and agent-level behavioral intelligence; subsequently, it establishes a theoretical framework for N-LMRMs. Empirical validation on state-of-the-art systems—including OpenAI’s O3 and O4-mini—demonstrates the framework’s effectiveness.
This study investigates how vision-language models integrate visual and textual information during chain-of-thought (CoT) reasoning and examines their susceptibility to misleading textual cues. Through dynamic tracking of CoT confidence trajectories, controlled interventions introducing deceptive text, and comparative analysis across 18 models, the work reveals a previously undocumented “answer inertia” phenomenon: models exhibit strong reliance on textual cues even when capable of correction. The findings indicate that while reasoning-focused training enhances corrective capacity, it fails to eliminate textual bias. Although instruction-tuned models less frequently cite misleading information explicitly, their reasoning traces more readily expose inconsistencies between visual and textual inputs. These results suggest that although CoT partially reflects multimodal integration, its apparent fluency may mask an implicit overreliance on textual signals.
Multimodal large language models (MLLMs) suffer from visual hallucinations and over-reliance on textual priors during visual reasoning. Method: We propose a tool-augmented agent architecture that decouples the LLM from a lightweight, specialized vision module, enabling fine-grained visual analysis and iterative reasoning via chain-of-thought guidance. We introduce a three-stage diagnostic evaluation framework to systematically uncover failure modes of mainstream MLLMs and design a modular, interpretable agent workflow supporting dynamic visual tool invocation and result verification. Contribution/Results: Our approach achieves +10.3 and +6.0 absolute improvements on MMMU and MathVista, respectively—surpassing same-parameter-scale models and approaching the performance of significantly larger ones. The code and evaluation framework are publicly released.
The effectiveness boundaries of multimodal Chain-of-Thought (CoT) reasoning remain unclear. This study systematically evaluates 12 tasks, comparing 14 non-reasoning and 8 reasoning-based multimodal large language models. It reveals that CoT enhances performance in mathematical, scientific, and multi-image reasoning tasks, yet degrades accuracy in perceptual tasks such as visual grounding and counting. The analysis identifies a pervasive “vision-light, reasoning-heavy” tendency—where visual information is progressively downweighted during reasoning—as a key bottleneck in multimodal CoT. Furthermore, the findings indicate that current open-source models exhibit limited overall improvement on such reasoning tasks, highlighting a critical gap in effectively integrating perception with structured reasoning.
Existing text-based chain-of-thought (CoT) reasoning treats visual input as static context, creating a “semantic gap” between perception and symbolic reasoning. To bridge this gap, we propose a paradigm shift—from “thinking about images” to “thinking with images”—and introduce the first systematic three-stage framework wherein vision serves as a dynamic cognitive workspace: visual input → manipulable intermediate representation → autonomous generation and operation. Our method integrates programmable visual operations, intrinsic imagination mechanisms, cross-modal alignment, and textual CoT, elevating vision from passive input to an active medium for reasoning. This framework establishes the theoretical foundation for “thinking with images,” advances multimodal AI toward human-like cognitive autonomy, and provides a clear roadmap for evaluation design, key technology validation, and future research directions.
This work addresses the opacity of reasoning processes in multimodal large language models (MLLMs) by identifying a novel diagnostic failure mode—“modality disruption”: high-confidence unimodal errors dominate multimodal fusion decisions, thereby suppressing corroborative evidence from other modalities. To diagnose this phenomenon, we propose a lightweight, model-agnostic framework that treats each modality as an independent agent. The framework integrates candidate label generation, self-evaluation prompting, and aggregation-based fusion to enable interpretable auditing of both modality-specific contributions and disruptive behaviors. Crucially, it is the first method to systematically characterize the dynamics of multimodal conflict. Evaluated on sentiment analysis benchmarks, it effectively disentangles data bias from intrinsic model deficiencies. Our approach provides a transferable, principled diagnostic tool for assessing the reasoning reliability of MLLMs.
This work addresses the challenge of modality isolation in complex, interleaved multimodal reasoning, where alternating text and image generation often leads to contextual drift in images and underutilization of visual information in text, thereby undermining cross-modal synergy. To mitigate this, the authors propose the MoTiF framework, which decomposes reasoning into atomic operations and introduces modality transition fidelity as a novel training signal. By quantifying cross-modal hallucination and insufficient visual grounding through a dedicated modality transition loss, MoTiF integrates reflective supervised fine-tuning with process-based GRPO reinforcement learning. Crucially, it enforces structured supervision at modality boundaries rather than relying solely on final-task accuracy. Experiments demonstrate significant improvements in both cross-modal consistency and task performance across four visual puzzle benchmarks, highlighting the critical role of transition-level supervision in interleaved reasoning.
Existing evaluation benchmarks struggle to diagnose the perception-reasoning disconnect in multimodal large language models (MLLMs) arising from insufficient visual evidence in complex urban scenes. To address this, this work introduces AD2-Bench, a novel benchmark, and EGVOR, an evidence-grounded visual reasoning framework that pioneers a hierarchical visual diagnostic mechanism. EGVOR explicitly models reasoning as spatial-semantic triplets—termed evidence atoms—and integrates structured evidence generation, hierarchical curriculum learning, and reflective reinforcement training to probabilistically uncover two core failure sources: spatial ambiguity and semantic uncertainty. Experiments demonstrate that EGVOR substantially enhances reasoning stability and accuracy under adverse conditions, offering an interpretable and diagnosable paradigm for trustworthy multimodal cognition.
Existing multimodal reasoning approaches often suffer from the loss of fine-grained visual details or compromised visual faithfulness due to suboptimal timing and manner of visual evidence integration. To address this limitation, this work proposes the CSMR framework, which introduces a novel cognitive scheduling mechanism: a language model dynamically controls an independent visual perception module, invoking task-relevant visual evidence on demand during reasoning. This design overcomes the constraints of static fusion and end-to-end joint optimization paradigms. Evaluated under zero-shot settings, the proposed method achieves substantial performance gains over state-of-the-art baselines across multiple multimodal benchmarks, demonstrating that dynamic scheduling effectively enhances both reasoning accuracy and visual faithfulness.
Existing multimodal benchmarks emphasize the fluency of chain-of-thought (CoT) generation in vision-language models but neglect whether such reasoning is genuinely grounded in visual evidence and logically coherent. Method: We introduce MM-CoT, the first diagnostic benchmark for multimodal CoT, jointly evaluating visual grounding and logical coherence through a structured event-chain selection task. It incorporates visual consistency verification, causal reasoning, and commonsense judgment, augmented by orthogonal constraints and adversarial perturbations to disentangle grounding failures from logical inconsistencies. Contribution/Results: Experiments reveal that state-of-the-art vision-language models perform substantially below human levels on MM-CoT, exposing a critical gap between generative fluency and reasoning fidelity. Moreover, MM-CoT exhibits low correlation with existing benchmarks, confirming its unique ability to quantify the long-overlooked dimension of reasoning reliability.
Multimodal large language models often suffer from inefficiency, verbosity, and hallucination due to their reliance on end-to-end generation or explicit linguistic reasoning chains. To address these limitations, this work proposes HIVE, a novel framework that recursively extends Transformer modules within an aligned latent space, enabling multi-step implicit “slow thinking” reasoning without requiring explicit textual justifications. HIVE injects hierarchical visual cues—from global scenes to fine-grained regions—into the latent reasoning process, thereby performing grounded, iterative inference while eschewing dependence on superficial language explanations. Experimental results demonstrate that incorporating hierarchical visual knowledge at test time significantly enhances complex scene understanding and overall model performance.