From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

πŸ“… 2026-09-28
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the inefficiency and error-proneness of vision-language models during multi-step compositional reasoning by proposing an internal state-aware adaptive early stopping strategy. Through controlled experiments involving hidden state probing and counterfactual image interventions, we uncover an internal β€œanswer-ready” mechanism within these models. To our knowledge, this is the first work to correlate such internal readiness with computational allocation during inference, enabling the training of a lightweight detector for dynamic early termination. Evaluated on benchmarks including MMStar, the proposed method reduces inference tokens by approximately 75% while improving accuracy by over three percentage points. These results demonstrate that dynamically adapting computation based on internal model states achieves simultaneous optimization of both reasoning efficiency and performance.
πŸ“ Abstract
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterfactual image pairs isolate changes in the visual evidence needed to answer. Across four models, direct answers, hidden-state readouts, and state interventions show that the individual judgments can be made without explicit reasoning and that intervening on the corresponding states can affect the answer. During reasoning, the Composite answer becomes decodable from hidden states and usable from shortened traces, often before the model stops on its own. We train a small detector to predict this readiness and stop reasoning at that point. On MMStar and RealWorldQA, this reduces mean reasoning tokens by 79.1% and 74.5%, while average accuracy rises by 3.13 and 3.30 percentage points, respectively. These findings connect the internal development of answer readiness to a practical rule for allocating reasoning computation.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Visual Reasoning
Composite Tasks
Internal Dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Reasoning Dynamics
Hidden State Intervention
Early Stopping
Answer Readiness
πŸ”Ž Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13