Distilling Visual Reasoning into Text Space

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of large vision-language models in reasoning tasks that transcend visible concepts, where intermediate visual representations often introduce errors and disrupt textual reasoning. To this end, we propose V2T, a novel chain-of-thought distillation paradigm from vision to text. Specifically, we first train an interleaved vision-text chain-of-thought teacher model, then integrate logits distillation, attention mapping, and reinforcement learning guidance to internalize visual reasoning capabilities into the student model's purely textual space, thereby eliminating reliance on intermediate image generation. Experimental results demonstrate that our method achieves an average accuracy improvement of 14.3% across multiple benchmarks, with training speed 42 times faster than existing approaches, significantly outperforming both the teacher model and baseline methods.
📝 Abstract
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introduce errors and increasingly interfere with textual reasoning as reasoning progresses. We propose Visual-to-Text Chain-of-Thought Distillation (V2T), a framework that enables LVLMs to internalize visual reasoning without generating intermediate visual representations at inference time. V2T first trains a teacher LVLM using interleaved visual and textual chains of thought, and then uses knowledge distillation to train a student LVLM using the teacher's logits and cross-entropy supervision from ground-truth textual reasoning. When reasoning images can be mapped to the original image, V2T can additionally distill the teacher's attention to corresponding regions, while ground-truth bounding boxes can further guide a subsequent reinforcement learning stage. Experiments across multiple multimodal reasoning benchmarks show that V2T consistently outperforms the teacher and existing baselines, improving average accuracy by 14.3% on a held-out set and 2.7% on the broader visual evaluation suite. Moreover, lightweight SFT and substantially reduced RL make V2T up to 42x faster to train than state-of-the-art baselines.
Problem

Research questions and friction points this paper is trying to address.

Large Vision-Language Models
Multimodal Reasoning
Visual Reasoning
Intermediate Visual Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual-to-Text Distillation
Chain-of-Thought
Knowledge Distillation
Large Vision-Language Models
Reinforcement Learning
🔎 Similar Papers