🤖 AI Summary
This study addresses the substantial computational overhead of visual tokens in multimodal large language models, noting that existing pruning methods permanently discard potentially useful information. To overcome this limitation, this work proposes δ-Vision, which reveals through low-rank intervention analysis that visual influence is concentrated within a low-dimensional subspace. Accordingly, it leverages lightweight MLPs to construct layer-wise visual memory that approximates visual states at each layer, directly replacing iterative Transformer computations. This approach enables efficient feature reconstruction and attention optimization while retaining all visual tokens. Extensive evaluations on image and video benchmarks demonstrate that δ-Vision surpasses pruning baselines in accuracy, significantly reduces computational costs, and achieves highly competitive inference efficiency.
📝 Abstract
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $\delta$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $\delta$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.