VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational redundancy in multimodal large language models caused by fixed dense visual encoding, which struggles to accommodate unevenly distributed visual information. We propose a pioneering native elastic visual representation framework that enables end-to-end adaptive visual granularity allocation via a gated spatial pooler and a granularity router. By integrating shared MRoPE coordinates with a self-distillation strategy, we conduct large-scale training based on the Qwen series, overcoming limitations of existing methods in content adaptivity, task generalization, and inference infrastructure integration. Experiments demonstrate that our approach reduces visual tokens by 43% on average while retaining 98.9% of performance. Furthermore, deployment with SGLang yields a 2.3× throughput improvement and a 54.4% reduction in time-to-first-token latency.
📝 Abstract
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Visual Token Efficiency
Elastic Visual Representation
Token Pruning
Computational Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Elastic Visual Representation
Multimodal Large Language Models
Gated Spatial Pooler
Granularity Router
Token Efficiency
💼 Related Jobs
No related jobs found.
Y
Yuan Feng
University of Science and Technology of China
Qize Yang
Qize Yang
Tongyi Lab, Alibaba Group
Computer VisionDeep Learning
Ruizhe Chen
Ruizhe Chen
Zhejiang University
LLMMLLM
Sibo Song
Sibo Song
Alibaba
computer visiondeep learningmultimodal learning
H
Haolin He
The Chinese University of Hong Kong
Muzhi Zhu
Muzhi Zhu
Zhejiang University
Computer VisionMachine Learning
Z
Zihan Liu
Alibaba Token Hub, Alibaba Group
Yunfei Chu
Yunfei Chu
Alibaba Group
machine learning
X
Xize Cheng
Alibaba Token Hub, Alibaba Group
Y
Yuxuan Wang
Alibaba Token Hub, Alibaba Group
Jin Xu
Jin Xu
Qwen Team, Alibaba Group
Multimodal InteractionLarge Language ModelSpeech SynthesisVideo/Audio Processing
X
Xike Xie
University of Science and Technology of China