SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-reliance of multimodal large language models on linguistic priors, as well as the high computational cost and lack of structural control in text-based chain-of-thought reasoning. We propose a Structured Latent Visual Reasoning framework built upon Qwen2.5-VL-7B. To our knowledge, this work is the first to establish an explicit functional structure for latent reasoning, decoupling it into typed stages such as planning and localization. By employing masked contrastive training and stage-wise supervision instead of autoregressive text generation, our approach enables fine-grained visual reasoning without explicit textual outputs. Experimental results demonstrate that the proposed method improves performance by 9.4% on MMVP and 14.2% on BLINK Relation, significantly outperforming baselines across multiple benchmarks. Overall, this framework effectively balances reasoning controllability with computational efficiency.
📝 Abstract
Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration. SLVR first trains the model to rely on the image by masking answer-revealing text and contrasting the correct answer with visually plausible distractors. It then organizes reasoning into latent stages for planning, grounding, evidence selection, and integration, supervising each stage with the corresponding signal: plans, boxes, visual evidence, and final rationales. This gives latent reasoning an explicit functional structure while avoiding generated textual chains at inference time. Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks, with absolute gains of +9.4 on MMVP and +14.2 on BLINK Relation, as well as improvements on V*, MathVista, and ChartQA. These results suggest that structured latent supervision can improve fine-grained visual reasoning without the decoding overhead of textual CoT. Project page is available \href{https://bogao-code.github.io/SLVR/}{here}.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Reasoning
Latent Reasoning
Chain-of-Thought
Linguistic Priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Latent Reasoning
Multimodal Large Language Models
Visual Reasoning
Chain-of-Thought
Contrastive Learning