🤖 AI Summary
This study addresses the computational redundancy in multimodal large language models caused by fixed dense visual encoding, which struggles to accommodate unevenly distributed visual information. We propose a pioneering native elastic visual representation framework that enables end-to-end adaptive visual granularity allocation via a gated spatial pooler and a granularity router. By integrating shared MRoPE coordinates with a self-distillation strategy, we conduct large-scale training based on the Qwen series, overcoming limitations of existing methods in content adaptivity, task generalization, and inference infrastructure integration. Experiments demonstrate that our approach reduces visual tokens by 43% on average while retaining 98.9% of performance. Furthermore, deployment with SGLang yields a 2.3× throughput improvement and a 54.4% reduction in time-to-first-token latency.
📝 Abstract
Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Downsampling sacrifices this detail, while existing token pruning and adaptive approaches remain limited in content-adaptive granularity, task generalization, and integration with modern MLLMs and serving infrastructure. Overcoming these limitations calls for foundation models that learn, end to end, where-and at what granularity-to allocate visual representations, a native capability we term elastic visual representation weaving. We introduce VisionWeave, establishing this capability in frontier-level MLLMs through large-scale training. It combines two components: a gated spatial pooler constructs coarse-grained representations alongside native fine-grained representations within a shared MRoPE coordinate, while a granularity router learns their content-adaptive allocation. Through self-distillation alone, we validate this capability on Qwen3.5-4B and scale to Qwen3.8-27B with over 30K A100 GPU-hours. Based on Qwen3.8-27B, VisionWeave adaptively adjusts token savings to visual content, saving 43.0% tokens on average while retaining 98.9% native performance across eight benchmarks, versus only 88% performance preserved for token pruning baselines with a fixed 50% savings target. Extensive evaluations confirm robust efficiency-quality trade-offs across diverse tasks, resolutions and video frames. When deployed on SGLang serving engine, our method achieves a 2.3x throughput gain while reducing mean TTFT by 54.4% and mean TPOT by 60.6%. Together, we believe these results position elastic visual weaving as a promising capability for next-generation multimodal models.