🤖 AI Summary
This study addresses the "symbolic-visual gap" in multimodal large language models, wherein perception and reasoning are misaligned, hindering precise evidence extraction from images for downstream inference. We are the first to identify this phenomenon, revealing that symbolic inputs can significantly guide attention toward correct visual evidence. Accordingly, we propose an annotation-free symbol-to-vision alignment strategy employing an online policy self-distillation framework. Through a residual token-level objective, the evidence selection capability inherent in the symbolic perspective is effectively transferred to image-input models. Extensive evaluations across diverse benchmarks and model scales demonstrate that our approach successfully bridges the bottleneck between perception and reasoning, yielding substantial improvements in multimodal capabilities, including chart document understanding, mathematical reasoning, and general visual question answering.
📝 Abstract
Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding the model, question, and decoding fixed, we replace images with their caption or code representations (symbolic views), which seems to be redundant given the clear image structures, but the performance surprisingly improves by 10.2% to 23.6% across model scales and datasets. We term this performance gap as the Symbolic Visual Gap and then take a closer look at it. Through experiments, we find that although the visual evidence can already appear in the reasoning trace for the image-input model, the symbolic-view-input model shows much higher attention to the correct evidence than the image-input model. This suggests that despite good capabilities from current works in perception and reasoning themselves, another bottleneck exists between perception and reasoning in selecting perceived visual information as appropriate evidence for subsequent reasoning. To handle this bottleneck, since the symbolic view steers attention toward correct evidence and is readily obtained at scale, it provides supervision for evidence selection without manually labeled evidence. Building on this, we introduce MM-OPD, a multimodal on-policy self-distillation framework for symbolic-to-visual correction that transfers guidance from symbolic-conditioned behavior to the image-conditioned policy through residual token-level targets, steering the model toward correct visual evidence. Experiments across benchmarks and model scales show that MM-OPD improves a broad range of multimodal abilities, with gains in visual perception, chart and document understanding, mathematical reasoning, and general VQA.