Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of language reasoning capabilities in multimodal large models (VLMs) during visual alignment. We propose LIFT, a method that transfers reasoning abilities from base large language models (LLMs) to VLMs via lightweight vector intervention. Specifically, LIFT defines reasoning vectors based on hidden state differences and demonstrates that vectors derived from LLMs significantly outperform those intrinsic to VLMs. Through vector injection and learnable adaptation, this capability transfer is achieved without retraining the backbone network. Experiments across six benchmarks show that LIFT effectively restores reasoning performance while revealing the positive influence of external vectors on intermediate reasoning behaviors.
📝 Abstract
Most Vision-Language Models (VLMs) are built by extending pretrained Large Language Models (LLMs) with visual modules and multimodal alignment. However, this multimodal scaling often degrades the language-side reasoning ability originally encoded in the base LLM. While the base LLM retains usable reasoning after scaling, the aligned VLM itself cannot reliably access this ability. Therefore, recovering the degraded reasoning capability in VLMs would benefit more from seeking help from the base LLM than from the VLM alone. Motivated by this, we propose LIFT (Language-side reasonIng Facilitation and Transfer), a lightweight vector-intervention method that transfers reasoning capability from the base LLM to the VLM without retraining the backbone. LIFT defines Reasoning Vectors as answer-token hidden-state differences between a Reasoner path with an explicit reasoning trace and a Solver path without it, and injects these vectors into language-side activations of the target VLM. LIFT further supports learnable vector adaptation while keeping the VLM backbone frozen. We evaluate LIFT on two VLMs across six reasoning benchmarks, comparing Reasoning Vectors extracted from the base LLM and from the aligned VLM under matched protocols. Results show that LLM-derived vectors consistently outperform VLM-derived vectors, confirming that the base LLM is a more effective source for recovering reasoning. LIFT partially recovers degraded reasoning through lightweight language-side interventions. Further analyses show that Reasoning Vectors influence intermediate reasoning behavior rather than merely altering final answers. The source code will be released soon.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Multimodal Reasoning
Reasoning Degradation
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reasoning Vectors
Vector Intervention
Vision-Language Models
Lightweight Adaptation
Multimodal Reasoning
Ziyi Wang
Ziyi Wang
University of Electronic Science and Technology of China
CVMLLMLLM
Li Li
Li Li
Southeast University
multi-modal
A
Aolin Zhou
School of Computer Science & Engineering, Southeast University, China
Y
Yankun Shen
School of Computer Science & Engineering, Southeast University, China
C
Chonghan Liu
University of California, Los Angeles, USA
S
Shuxia Lin
School of Computer Science & Engineering, Southeast University, China
X
Xu Yang
School of Computer Science & Engineering, Southeast University, China