Multimodal Latent Reasoning via Hierarchical Visual Cues Injection

📅 2026-02-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Multimodal large language models often suffer from inefficiency, verbosity, and hallucination due to their reliance on end-to-end generation or explicit linguistic reasoning chains. To address these limitations, this work proposes HIVE, a novel framework that recursively extends Transformer modules within an aligned latent space, enabling multi-step implicit “slow thinking” reasoning without requiring explicit textual justifications. HIVE injects hierarchical visual cues—from global scenes to fine-grained regions—into the latent reasoning process, thereby performing grounded, iterative inference while eschewing dependence on superficial language explanations. Experimental results demonstrate that incorporating hierarchical visual knowledge at test time significantly enhances complex scene understanding and overall model performance.

Technology Category

Computer Vision: Multi-modal VisionNatural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Multimodal Learning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
The advancement of multimodal large language models (MLLMs) has enabled impressive perception capabilities. However, their reasoning process often remains a"fast thinking"paradigm, reliant on end-to-end generation or explicit, language-centric chains of thought (CoT), which can be inefficient, verbose, and prone to hallucination. This work posits that robust reasoning should evolve within a latent space, integrating multimodal signals seamlessly. We propose multimodal latent reasoning via HIerarchical Visual cuEs injection (\emph{HIVE}), a novel framework that instills deliberate,"slow thinking"without depending on superficial textual rationales. Our method recursively extends transformer blocks, creating an internal loop for iterative reasoning refinement. Crucially, it injectively grounds this process with hierarchical visual cues from global scene context to fine-grained regional details directly into the model's latent representations. This enables the model to perform grounded, multi-step inference entirely in the aligned latent space. Extensive evaluations demonstrate that test-time scaling is effective when incorporating vision knowledge, and that integrating hierarchical information significantly enhances the model's understanding of complex scenes.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
latent space
visual cues
hallucination
chain of thought
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Latent Reasoning
Hierarchical Visual Cues
Slow Thinking
Latent Space Alignment
Iterative Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yiming Zhang
Nanyang Technological University
Q
Qiangyu Yan
Huawei Noah's Ark Lab
B
Borui Jiang
Huawei Noah's Ark Lab
K
Kai Han
Huawei Noah's Ark Lab