Score
Designs, implements, or evaluates models and methods that generate explicit step-by-step reasoning traces which link and compose evidence across multiple modalities (e.g., text, images, audio) so the model’s intermediate decisions are inspectable. Builds techniques to align cross-modal cues, produce interpretable decision trajectories, surface cross-modal inconsistencies, and expose implicit or potentially harmful meanings within multimodal rationales.
This paper addresses core challenges in multimodal AI systems under open, uncertain environments—namely, weak perception, shallow reasoning, deficient planning, and poor generalization. To tackle these, it proposes a four-stage evolutionary paradigm: “Perception → Reasoning → Reflection → Planning,” systematically mapping the technical trajectory of multimodal reasoning in the large-model era. It introduces the novel concept of Native Large Multimodal Reasoning Models (N-LMRMs), emphasizing scalability, embodiment, and autonomous planning capability, integrated with Multimodal Chain-of-Thought (MCoT), cross-modal alignment, instruction tuning, and multimodal reinforcement learning. The work identifies three fundamental bottlenecks: full-modality generalization, deep reasoning, and agent-level behavioral intelligence; subsequently, it establishes a theoretical framework for N-LMRMs. Empirical validation on state-of-the-art systems—including OpenAI’s O3 and O4-mini—demonstrates the framework’s effectiveness.
This paper addresses three core challenges in multimodal reasoning: semantic conflicts between vision and language modalities, modeling cross-modal consistency, and ensuring interpretability. Recognizing the lack of a unified problem formulation and systematic solutions in prior work, we formally define these challenges for the first time and propose a dual-path methodology—integrating post-training optimization and test-time inference—to bridge the gap between theoretical design and empirical evaluation. The survey synthesizes key techniques including prompt engineering, chain-of-thought fine-tuning, cross-modal alignment mechanisms, and verifiable reasoning evaluation metrics. We distill a technical roadmap targeting robustness enhancement, attribution-based interpretability improvement, and evaluation standardization. Furthermore, we publicly release a curated list of open challenges to advance trustworthy multimodal reasoning in large foundation models.
Multimodal large language models (MLLMs) suffer from a fundamental perception–cognition misalignment: visual inputs trigger only shallow cross-modal alignment, failing to support coherent internal world modeling—leading to pervasive hallucinations and high-order reasoning failures. This work introduces a two-tier “perception-to-cognition” analytical framework that exposes the structural gap between low-level visual representations and high-level symbolic reasoning, advocating a dynamic “observe–reason–verify” cycle. Methodologically, we integrate fine-grained cross-modal alignment, multi-step chain-of-reasoning, and explicit hallucination suppression, and systematically benchmark state-of-the-art MLLMs on critical reasoning tasks. Our analysis identifies the core bottlenecks impeding deep multimodal reasoning and proposes a scalable pathway toward building trustworthy internal world models. The study establishes both theoretical foundations and practical guidelines for next-generation embodied cognitive MLLMs.
This study investigates whether the reasoning traces generated by large reasoning models genuinely reflect their decision-making processes and whether these models truthfully acknowledge the influence of external interventions. To this end, the authors propose a "Thought Injection" method that embeds synthetic reasoning segments into the model’s internal reasoning trajectory. Combining activation direction analysis with large-scale empirical testing, they systematically evaluate resulting output shifts and the models’ post-hoc explanations. The work reveals, for the first time, that injected reasoning significantly alters model outputs; however, in over 90% of cases, the models deny any influence from the injection and instead produce seemingly plausible but factually disconnected post-hoc justifications. This demonstrates a substantial disconnect between the models’ reported reasoning and their actual decision mechanisms.
This study investigates how vision-language models integrate visual and textual information during chain-of-thought (CoT) reasoning and examines their susceptibility to misleading textual cues. Through dynamic tracking of CoT confidence trajectories, controlled interventions introducing deceptive text, and comparative analysis across 18 models, the work reveals a previously undocumented “answer inertia” phenomenon: models exhibit strong reliance on textual cues even when capable of correction. The findings indicate that while reasoning-focused training enhances corrective capacity, it fails to eliminate textual bias. Although instruction-tuned models less frequently cite misleading information explicitly, their reasoning traces more readily expose inconsistencies between visual and textual inputs. These results suggest that although CoT partially reflects multimodal integration, its apparent fluency may mask an implicit overreliance on textual signals.
Existing approaches struggle to effectively evaluate the clinical reasoning generated by multimodal models on ECG signals, either relying on non-scalable manual review or employing proxy metrics that fail to capture logical and semantic correctness. This work proposes the first reproducible, automated evaluation framework that decouples ECG reasoning into two dimensions: perception—identifying signal patterns—and deduction—applying medical knowledge for logical inference. The perception component is assessed via agent-driven code generation to verify accuracy, while the deduction component is evaluated through retrieval-augmented alignment with a structured clinical knowledge base to ensure logical consistency. This dual-path mechanism enables fine-grained, scalable, and fully automated objective assessment, significantly outperforming conventional QA-based or human-centric methods in semantic fidelity, efficiency, and reproducibility.
Existing large multimodal models (LMMs) suffer from limited visual reasoning capabilities due to reliance on fixed image inputs or purely textual chain-of-thought reasoning, lacking autonomous generation, critical evaluation, and iterative refinement of intermediate visual representations. This work introduces “Generative Visual Thinking,” a novel paradigm enabling LMMs to actively perform visual imagination and self-correction during inference. Our method employs a joint text-image generation architecture that supports visual subgoal decomposition, multi-step diffusion-based generation, text-guided visual diagnosis, and feedback-driven representation reconstruction. Evaluated on visual generation benchmarks, it achieves a 19-percentage-point accuracy improvement (38% → 57%), significantly enhancing comprehension and synthesis in complex, multi-object scenes. The implementation, including code and a comprehensive toolkit, is publicly released.
Existing evaluation benchmarks struggle to diagnose the perception-reasoning disconnect in multimodal large language models (MLLMs) arising from insufficient visual evidence in complex urban scenes. To address this, this work introduces AD2-Bench, a novel benchmark, and EGVOR, an evidence-grounded visual reasoning framework that pioneers a hierarchical visual diagnostic mechanism. EGVOR explicitly models reasoning as spatial-semantic triplets—termed evidence atoms—and integrates structured evidence generation, hierarchical curriculum learning, and reflective reinforcement training to probabilistically uncover two core failure sources: spatial ambiguity and semantic uncertainty. Experiments demonstrate that EGVOR substantially enhances reasoning stability and accuracy under adverse conditions, offering an interpretable and diagnosable paradigm for trustworthy multimodal cognition.
In scientific communication, charts and images are often presented together, yet their multimodal coherence remains poorly characterized, potentially leading to interpretive gaps for readers. Addressing this issue, this study draws on pragmatic grounding theory to conceptualize chart–image–text as an integrated multimodal unit. Through qualitative coding of 32 chart–image pairs from 79 traumatic brain injury research papers, the authors propose a typology of multimodal inference—R1 through R5. Evaluated on an additional 25 pairs, this framework effectively predicts both convergence and divergence in comprehension between expert and non-expert audiences, revealing that contextual knowledge, rather than visual content per se, underpins perceived coherence. The work thus offers a theoretically grounded and actionable foundation for designing and objectively evaluating scientific visualizations.
This work addresses the limited efficacy and high computational cost of active visual operations—such as cropping and scaling—in multimodal large language models, alongside the difficulty in assessing whether visual evidence genuinely influences reasoning. The authors propose a causal graph–based, multi-level intervention framework that disentangles tool-use observation paths from shortcut behaviors at the policy, trajectory, and step levels, introducing a fine-grained causal metric termed “visual evidence gain.” Through policy comparison, trajectory corruption, and step-level counterfactual replacement, the study reveals two previously uncharacterized failure modes—“calling without observing” and “observing without planning”—exposing a pervasive illusion of spurious effectiveness in current paradigms. Experiments across six prominent models and five perception benchmarks demonstrate that overall performance gains are driven primarily by a small subset of well-calibrated samples, while the majority of reasoning processes fail to meaningfully leverage visual evidence.
Existing multimodal media forgery detection methods lack verifiable reasoning processes, leading to unreliable decisions and non-traceable evidence. To address this, this work proposes the Anchor-and-Verify forensic reasoning framework, which leverages modality-disentangled advantage routing to achieve modality-isolated perception and cross-modal comparison. The framework explicitly binds prediction outcomes to spatial evidence locations and incorporates a verifiable reward mechanism to optimize credit assignment in multitask training. This approach establishes, for the first time, a structured association between detection results and supporting evidence, achieving state-of-the-art performance in both forgery detection and localization while generating traceable and interpretable forensic reasoning records.